Multi-model image platforms sell optionality. Most desks buy confusion. A dropdown with Nano Banana, GPT Image 2, and a video-capable roster looks like coverage until nobody can say which failure each option is supposed to absorb. Before a team debates taste, it needs a fixed battery of ugly inputs. That is the practical reason to trial AI Imagine: not because the menu is long, but because one workspace can run the same dirty stills without resetting accounts between chores.
Data-minded buyers already know the pattern from model registries. Average demo quality is a weak KPI. Hard-case recall is the number that predicts whether a tool survives a real week. The evaluation below treats generative image work like a small ML pilot: freeze the inputs, name the failure labels, then read keep-rate against credit spend.
Model Menus Hide A Missing Evaluation File
Teams often switch models when a result feels “off,” then call the switch an insight. Without a frozen input set, that behavior is fashion. The second model may look better only because the prompt drifted, the crop changed, or the reviewer was tired. An evaluation that cannot be rerun tomorrow is not an evaluation.
The missing artifact is usually a one-page battery: three source files, three reject labels, and a place to record whether an export would enter the shared folder. Until that page exists, procurement meetings argue about brand names while the production queue keeps regenerating the same soft failures. A week of “try the new model” notes in chat will not survive staffing changes; a dated scorecard will.
Average Pretty Is Not A Keep Metric
Pretty is what demos optimize. Keep-rate is what budgets feel. Define keep-rate as exports that would clear the desk without another generate. If four of six runs look impressive in isolation but only one would ship beside the source, the keep-rate is 1/6, not “mostly good.” That fraction is what later justifies Standard versus Pro credit burn.
Assemble Three Dirty Stills Before Any Prompt
Clean product shots flatter every model. Dirty stills expose the seams. Build the battery once and refuse to change it mid-comparison. Ideal files are already in most marketing drives: a motion-blurred event photo, a social screenshot with a platform watermark, and a title card dense with small lettering. None of these should be confidential customer data.
Label the jobs in plain language before opening the tool. Blur case: recover a usable hero without inventing a new face. Watermark case: clear the mark without melting nearby edges. Title-card case: keep lettering readable after a background cleanup. If a request needs art direction, park it outside the battery. Mixing redesign with repair is how scorecards become vibes.
Blur Watermark And Title Card Cases
Write the three prompts as chores with an explicit “do not change identity” clause where faces appear. Keep aspect and resolution choices boring and fixed for the whole pilot—for example, stay on the workspace’s Auto/1K path if that is what daily users will click. The point is comparable friction, not a trophy render.

Workflow Steps That Keep The Battery Comparable
AI Imagine’s published image path is short enough to script: upload, describe the edit, generate, download. Use that path for every cell in the battery so the only intentional variable is the model or quality tier under test. Do not bounce to a separate upscaler for one cell and not the others. Side tools make the scorecard about glue, not about the platform.
Record wall-clock only as a secondary note. Primary columns should be keep / soft reject / hard reject, plus a one-line defect. Soft reject means salvageable with a tighter instruction. Hard reject means the export invents structure, warps edges, or leaves lettering unreadable next to the original. If two reviewers disagree on keep versus soft reject, mark soft reject and tighten the prompt once—do not average their opinions into a fake pass.
- Freeze the three source stills and store checksums or unchanged filenames.
- Run each still once on the Standard-cost path, then once on the Pro-cost path if the first keep-rate is weak.
- Score each export against the source still at thumbnail size and at the intended crop.
- Stop after one retry per cell; endless regenerations hide the base rate.
Score Identity Lettering And Edge Integrity
Use a tiny table so two reviewers can disagree in public instead of in private Slack taste threads.
| Case | Keep signal | Hard reject |
| Blurred event still | Face and outfit stay recognisable after cleanup | New facial geometry or plastic skin that cannot publish |
| Watermarked screenshot | Mark gone; nearby UI edges remain coherent | Melted icons, warped bars, or leftover texture smear |
| Dense title card | Words stay readable at post size | Letters fuse, drop, or become decorative noise |
When lettering is the job, GPT Image 2’s stronger in-image text handling is the fair specialist lane. When identity preservation after local edits is the job, Nano Banana’s character-consistency and scene-preservation claims are the lane to pressure. The battery exists so those lanes are assigned by failure type, not by whichever name is trending.
Read Credit Cost Against Keep Rate
Pricing on AI Imagine makes the arithmetic concrete enough for a pilot ledger. Image generations start at 2 credits on the Standard model path and 6 credits on the Pro model path. Monthly subscription credits refresh each cycle and unused credits do not roll over; credit packs last a year. That means a desk that “explores” without a battery is not only wasting staff time—it is burning inventory that expires.
A simple pilot math check: six battery cells on Standard ≈ 12 credits; the same six on Pro ≈ 36 credits. If Standard already yields a keep-rate the desk can live with, spending Pro credits on every exploratory click is theater. If Standard keep-rate is near zero on watermark and title-card cells, Pro is not a luxury—it is the minimum path for those chores. Either way, the image generator decision becomes a measured keep-rate question instead of a vibe vote. Starter’s 100 monthly credits are enough for a careful pilot; they are not enough for unstructured browsing if unused credits expire at cycle end.
Where The Battery Still Leaves Gaps
A three-still battery will not certify video models, audio tools, or niche converters on the same site. It also will not settle rights questions for press or user-generated photos. Treat the scorecard as a gate for still repair and generation quality, then open separate pilots for motion or format utilities. Expanding the battery mid-quarter usually reintroduces fashion under a new filename.

Keep The Battery File Next To The Invoice
Use AI Imagine when the team can store three dirty stills, run Upload → Describe → Generate without side quests, and accept a keep-rate number before upgrading plans. Skip it as the first purchase when nobody will maintain the battery, or when the real need is a single designer already shipping clean files without generative help.
The durable output of the pilot is a short scorecard that survives staff turnover, not a favorite model name. Put that file next to the invoice. The next budget review then argues about measured failures, not about which logo looked exciting in a demo.