Playbook · How the studio scales beyond one person
The production playbook
A studio isn't a person who's good with tools — it's a documented system anyone trained can run. This is the playbook the demo pipeline implements, written to be handed over.
Workflow
Every job enters through the same door and leaves through the same gate.
- 01Intake
One product photo + minimum context
- 02BriefHuman
One structured brief drives every deliverable
- 03Route
Each asset to its cheapest sufficient model
- 04Generate
Stills, adaptations, motion, audio
- 05GateHuman
Fidelity · text · brand safety
- 06Deliver
Named, versioned, ready to ship
Two human checkpoints, deliberately placed: at the brief (before any spend) and at the gate (before delivery). Everything between is automated.
Intake takes one grounding asset — the product photo — plus the minimum viable context. The vision model extracts everything it can see and interviews only for what it can't. An intake form that fills itself is the difference between a pipeline teams adopt and one they route around.
One structured brief then drives every deliverable. That's deliberate: consistency across a pack comes from shared upstream context, not from post-hoc correction.
Prompt system
Prompts are templates, not incantations. Nothing depends on a hero prompt someone keeps in their head.
BRIEF mood · setting · palette · headline EN/FR
+
FORMAT BLOCK "Square 1:1 promo tile. Render "{headlineFR}"
in clean bold type, high contrast…"
+
STYLE SUFFIX "commercial retail quality, sharp focus"
=
PROMPT ────────────────▶ nano-banana-proEach deliverable has a prompt builder, not a prompt. Change one field in the brief and eight prompts update. That is what makes versioning — EN/FR, seasonal, per-format — a parameter instead of a project.
- Subject and action first; camera, light and setting second; style descriptor last.
- Real product attributes injected from vision analysis — never invented.
- Video prompts carry an explicit Audio: cue; native-audio models reward sound design written into the prompt.
- Negative prompts are maintained per style, with artifact patterns added as QA finds them.
- No brand names or logos in generation prompts — lockups composite in post, where they're controlled.
- For product motion, references beat a first frame: reference-to-video models hold identity across the take, addressed positionally as [Image1], [Video1] with a stated job each.
Model routing
Draft cheap, finish premium. Routing — not negotiation — is where an AI studio finds most of its cost efficiency.
| Job | Tier | Why |
|---|---|---|
| Format adaptations | Flash-tier stills | Hero already set the look |
| Text-in-image tiles | Pro-tier stills | Only tier that renders type reliably |
| Video drafts | Veo Fast / Kling | Iterate at a fraction of the cost |
| Hero finish | Veo Standard | Only where the quality ceiling is the point |
| Volume cutdowns | Kling / Seedance | 5–7× cheaper, consistency holds |
The live table lives on the model landscape page and is re-evaluated monthly.
Quality gates
Every asset clears three checks before it leaves the studio.
Product fidelity
Is this the actual product — shape, colors, packaging — or a plausible look-alike?
Grounding every generation on the real photo makes this passable. The gate makes it guaranteed.
Text integrity
Is every rendered word spelled and accented correctly, in both languages?
French is reviewed by a French speaker, never assumed. Generated prices are checked against the source of truth.
Brand & claim safety
Any implied claims, accidental third-party marks, or missing mandated elements?
Composition must leave room for price points and legal lines before it leaves the studio.
Cost governance
Nobody spends a credit blind.
Pre-flight estimate
Shown and confirmed before any live generation.
One price config
List prices live in one file; the estimator reads from it.
Visible spend
Session spend tracked and always on screen.
Free demo mode
The full pipeline mocks itself — UX work and training never burn credits.
Cheap defaults
Cost-efficient tiers are default; premium is an explicit choice.
Handles, not reruns
Assets generated with trim room so a near-miss is an edit, not a regeneration.
Guardrails
The generative-AI guidelines this work is held to, restated as decisions rather than clauses. They exist so the studio can move fast — the risky questions are answered before anyone is mid-render.
The standing permission
AI may be used as a creative and production tool wherever the result is original, properly authorised, non-impersonative, accurate, and not misleading to a consumer. Everything already required — Legal, Regulatory, Privacy, Procurement, InfoSec, Brand, advertising approval — still applies. AI is a new way to make the asset, not a new route around the approvals.
Three lines that do not move
Nobody real, unless they said yes
No cloning, imitating or deliberately evoking the voice, likeness or recognisable characteristics of an actual person without documented rights. Celebrities, influencers, performers, customers, employees, executives — the same line for all of them, and for third-party characters and branded voices. Original synthetic creation is the preferred route; replication is not a shortcut, it is the thing being prohibited.
The overall impression is the test
Nothing may materially misrepresent the product, its appearance, quantity, function or performance; a person's identity or relationship to the brand; a customer experience, testimonial or endorsement; or a relationship with another company. Judged the way a reasonable consumer would take the finished ad as a whole — not by whether each component survives inspection on its own. An ad can be assembled entirely from true parts and still fail this.
The product is not the part you may invent
Where an ad shows a specific item that is for sale, that item must originate from authentic capture of the real thing. AI may work on everything around it. It may not quietly become the thing itself.
On a real product, where the line actually falls
This is the most specific rule in the set, so it gets the most specific treatment. Everything on the left is production craft. Everything on the right changes what the customer thinks they are buying.
AI may
- Cleanup and retouching
- Background replacement or extension
- Removing production artefacts
- Reframing and outpainting
- Lighting and colour correction that stays faithful
- Environment and set dressing
AI may not
- Replacing the captured product with a synthetic one
- Adding toppings, fillings or inclusions that were not there
- Materially increasing apparent quantity or portion
- Improving texture, colour, doneness or quality
- Fabricating preparation results or functionality
- Materially altering packaging or product attributes
Depicting people
Synthetic people, by surface
Fully synthetic talent is available for stills — lifestyle, social, digital, print, display — provided it is an original creation and not a copy of an identifiable person. For OLV, television and broadcast, the conservative default holds: not without specific review and approval. That is a position about consumer acceptance and talent agreements rather than about capability, and it is expected to move.
Hands are not talent
An AI-generated hand entering frame to press a button is an incidental element, not a performer. It stays incidental while nobody is identifiable, it is not built from a real person's likeness, it is not acting as a spokesperson, and the action it performs is an honest representation of using the product.
Voices
Original synthetic voice is fine for radio and digital audio. It must not clone a real individual, imitate an identifiable person, impersonate a recognisable character or protected brand voice, or leave a listener wrong about who is speaking — and the commercial-use rights have to be real.
Manufactured authority
No synthesised customer reviews, celebrity endorsements, employee statements, expert opinions, regulated-professional recommendations or before-and-after stories. The failure here is not the pixels; it is inventing a person whose credibility is doing the selling. Openly fictional scenarios are a normal Legal question, not this one.
What has to survive an audit
- Approved tools only
- Platforms cleared under the applicable technology, security, procurement and privacy requirements. That is a list someone maintains, not a judgement call at 6pm.
- What never gets uploaded
- Confidential information, personal information, biometric source material, talent recordings, licensed content, or third-party assets whose AI-processing rights have not been confirmed — unless that specific system is authorised for it.
- Retain the real source
- Where AI materially assists final product imagery, the authentic source frame stays in the asset-management system and stays traceable to the delivered creative. If you cannot produce the original, you cannot defend the final.
- The record
- Platform or vendor, what the AI component actually was, the approved source asset, confirmation of commercial-use rights, talent consent where relevant, and any specific exception granted. Meaningful provenance — not a log of every routine retouch.
- Accountability does not transfer
- Marketing, agency and production partners own the finished asset whether or not AI touched it. There is no version of this where the model is responsible.
Two questions that settle most of it
Product fidelity
If someone bought this because of this ad, would what arrives reasonably match what they were shown?
Uncertain is a no. Escalate to Legal or Regulatory.
The catch-all
Could a reasonable consumer be materially misled about who or what they are seeing or hearing, what they are buying, what it does, or whether a real person took part or endorsed it?
Yes or unsure — stop and escalate before publication.
Where this is enforced rather than promised
A guideline nothing implements is a hope. These are the places the studio in this demo makes the rule structural — so following it is the default path, not the disciplined one.
- Prices are never invented
- Vision autofill returns an empty field rather than a guess when no price is legible on the pack, and the UI flags it as needing a human. A wrong price is a compliance failure, not a typo.
- Reconstructed angles are labelled
- Packshot views the model extrapolated rather than grounded in a supplied reference are marked for label QA, so nobody mistakes a plausible back-of-pack for a photographed one.
- The prompt ships with the asset
- Every generated output exposes the prompt that made it, which is the provenance record the guidelines ask for, produced as a by-product rather than as homework.
- Product identity is pinned, not hoped for
- Reference-to-video binds the pack to a supplied still and instructs against drift, because a product that morphs mid-shot fails the fidelity test even when nobody intended it to.
- The demo states its own limits
- Where the pipeline cannot guarantee something — music that is not frame-synced, a take that needs an alignment pass — it says so in the interface rather than in a footnote.
Two things that are not exceptions
Food, and anything aimed at children
Every existing food-advertising, regulatory and child-directed requirement applies exactly as before. Extra care where creative is child-directed, where children are synthesised, where child-oriented characters appear, or where the media buy is aimed at children. AI does not create an exemption from rules that already exist.
Saying so
Material AI use is disclosed internally, in the approval workflow, always. External disclosure is decided by context — channel, platform, contract, regulation. The one hard case: where staying quiet about AI would itself make the ad misleading, it goes to Legal before it goes out.
This page expires. These guidelines get reviewed against regulation, platform and broadcaster requirements, talent agreements, campaign learnings and the quality of synthetic media itself. The clause most likely to move first is the broadcast one — synthetic on-camera talent is a conservative default about consumer and legal acceptance, not a permanent technical judgement.
Pre-flight
The guardrails above, as something you run against an actual asset. Tick it before the asset ships, not after someone asks.
Before you generate
If a real product is shown
If a person appears
Before it ships
Anything left unticked is a question someone will ask later, when it is more expensive to answer.
Teaching
Built to be handed over, not held onto.
Every generated asset in the demo exposes its prompt — the pipeline shows its work. That's the teaching model: every workflow visible, every decision documented, every playbook written so the second person, and the tenth, can run it.
Capability that lives in one person's head isn't a studio. It's a bottleneck with a title.
Measurement
What the studio is accountable for, and what good looks like.
| Measure | Tracked as | Good looks like |
|---|---|---|
| Speed | Brief-to-delivery time per pack | Same-day for standard versioning |
| Volume | Assets delivered per week, and per dollar, by type | Up and to the right, per dollar flat |
| Quality | First-pass approval rate through the gates | Rising; rework rate falling |
| Adoption | Teams and brand partners actively using the studio | Repeat intake, not one-off curiosity |
| In-housing | Share of work absorbed from external vendors | Where quality and rights make sense |
| Cost | Cost per finished asset vs. traditional baseline | Order-of-magnitude, not percentages |