1. Why every agent matters
Story generation is not one call — it is a chain. Each stage feeds the next, so a weakness anywhere ripples forward:
| Stage | If it is weak, you see |
|---|---|
| Concept Architect | Generic premise, flat characters, loose arc — everything after feels unfocused |
| Story Scriptwriter | Repetitive beats, vague scene descriptions, dialogue that sounds the same for every character |
| Effects Director | Overused camera gimmicks or none at all; pacing feels rushed or dragging |
| Asset Discovery | Missing props/locations, inconsistent character states, visuals that do not match the script |
Because later stages inherit earlier outputs, fixing the script rarely rescues a weak concept. Review the chain from the top when a story feels off — the root cause is often an earlier agent, not the one where the symptom appears.
Basics are covered in Agents — Customize, Swap and Earn Credits and Story Generation.
2. Precise instructions, low-quality model
This is the most common mismatch. The instruction is detailed and careful, but the model underneath cannot fully honor it:
What “precise” looks like
- Step-by-step structure, length limits, tone rules, examples
- Demands like “keep continuity across 12 chapters” or “vary dialogue by personality”
- Formatting constraints for the pipeline to parse reliably
What a lighter model may do
- Follow the first half of the instruction and drift by the second half
- Simplify characters or repeat phrasing instead of inventing fresh turns
- Miss subtle constraints like “15 seconds max per shot” or “states must cover 0 to duration”
Typical symptom: you spent time crafting a great prompt, but generations feel “almost right” — correct shape, weak execution, extra cleanup in the Storyboard tab. The instruction is not the problem; the model’s capacity is.
The opposite mismatch also happens — a strong model with a vague, generic instruction. Then you pay a higher per-generation cost without getting the distinctiveness the model can deliver. Precision and model strength need to move together.
3. Balancing instructions and model quality
Think of an agent as direction + ability. Good results come when both are matched. Use the picker’s price per unit as a rough proxy for ability, and your instruction detail as direction:
| Combination | Cost | When to use |
|---|---|---|
| Precise + strong model | Higher | Final runs, marketplace agents, stories where voice and structure matter most |
| Precise + lighter model | Lower | Quick drafts to test structure; expect to polish or re-run with a stronger model |
| Simple + strong model | Higher but wasteful | Rarely ideal — add at least tone and constraints to get value |
| Simple + lighter model | Lowest | Throwaway experiments and chapter-count tests |
Practical balancing rules
- Put strength where writing matters. Use your best model for Concept Architect and Story Scriptwriter. Effects Director and Asset Discovery can often use a lighter, cheaper model without visible loss.
- Match complexity to model. If your Scriptwriter demands nuanced dialogue per character and tight second-by-second timing, pair it with a model that handles long-form coherence. A lighter model will flatten those nuances.
- Draft cheap, polish strong. Generate a short 4–6 chapter draft cheaply to validate the arc, then swap in a stronger Scriptwriter for the full run.
- Test before committing. Each agent card has a Test area with stage-specific sample input. Run it — a real generation at the selected model’s rate — to see if the instruction survives execution.
Costs are per text generated and shown in the picker. Keep an eye on the Dashboard’s credit chart — see Startup Credits — to spot which stage dominates spend when you mix models.
4. Web search — some agents can, some cannot
Models differ in tool support. Some support live web search or browsing, others are purely knowledge-based with a fixed cutoff date. This matters when freshness matters:
When search helps
- Stories needing current facts — recent events, real places, evolving tech
- Concept research that benefits from up-to-date references
- Verification of niche details you want grounded accurately
When it does not
- Invented worlds, fantasy, timeless drama — cutoff rarely matters
- Voice-driven dialogue and stylistic work — creativity over lookup
- Tight storyboard pacing — extra tool calls can slow generation
| Model capability | What it means for stories | Watch out |
|---|---|---|
| Supports web search | Can pull fresh information when the instruction asks for it | May be slower or slightly more expensive; needs clear guidance on when to search |
| No web search | Faster, often cheaper; relies on training data cutoff | May hallucinate recent facts — avoid for “set in 2026 with real news” concepts |
How to handle marketplace agents: the card shows model and description but not the prompt. If freshness matters, prefer agents whose description mentions “web-search aware” or “current events”, and test with a prompt that needs a recent fact. If the answer cites a pre-cutoff event as current, that model likely cannot browse — swap to one that can for that stage.
- You can mix. Use a search-capable model for Concept Architect where real-world grounding helps, and a non-search, strong writer for Scriptwriter where style matters more. Each stage’s model is independent.
- Be explicit in the prompt. If the model can search, tell the agent when to search (“look up X if the story is set after 2024”) and when not to — otherwise it may search unnecessarily or not at all.
- Ask the app: the model picker and agent editor reflect each model’s current capabilities. When in doubt, test the same instruction on two models — one with search, one without — and compare freshness vs. speed.
5. Diagnosing which side is the problem
When a generation disappoints, isolate:
| Symptom | Likely cause | Try |
|---|---|---|
| Follows half the instruction, ignores edge constraints | Model too light for the instruction’s complexity | Keep the prompt, swap to a stronger model and re-test |
| Generic output despite strong model | Prompt too vague | Add tone, examples, and structure; clone the built-in as a baseline |
| Facts are dated or hallucinated | Model without web search used for fresh-topic story | Switch that stage to a search-capable model or remove real-world dependency |
| Good concept, weak dialogue | Concept strong, Scriptwriter mismatched | Upgrade only the Scriptwriter stage — no need to change the whole pipeline |
6. Practical fixes and presets
- Clone before tuning. Use Clone on the Agents page to start from the built-in — you keep version history and can compare side-by-side.
- One stage at a time. Change only Concept Architect, re-generate, assess. Then move to Scriptwriter. Bulk swaps hide which change helped.
- Use Settings for signature, overrides for experiments. Save your best balanced set as global defaults in Settings → Default AI Agents; use Dashboard → New Idea → Custom AI Agents for single-story tests.
- Leverage marketplace wisely. A well-rated Scriptwriter with a strong model can rescue dialogue without you prompt-engineering from scratch. Check ratings — only verified buyers can rate.
- Mind archiving. Archived agents disappear from selectors; if a favorite vanishes, check My Agents filters.
Cost tip: marketplace purchases and tests both cost credits (see credits and billing). Draft with lighter models, then spend where quality shows most — usually Concept and Scriptwriter.
7. Pre-flight checklist
- Does every stage use an agent whose instruction detail matches its model strength?
- For fresh-topic stories, is at least Concept Architect on a search-capable model?
- Have you run Test on each customized agent with its stage’s sample input?
- Are your global defaults set, so New Idea pre-fills correctly?
- Check the picker price vs. expected length — a long 20-chapter story on a premium model will cost noticeably more.
If results still feel off, try swapping just one stage’s model, re-run, and compare. Small, isolated changes teach you the balance faster than rewriting everything at once.
