Back to Briefing

Same AI Chaos or Compounding Value

Same AI Chaos or Compounding Value

In This Issue

  • Fluent output is no longer the AI bottleneck. Your operating system around it is.

  • Five practices that keep AI producing value instead of volume, from a live operation

  • Labs receipt: the linter that made most of my editorial review deterministic

  • One CEO-survey number that explains the enterprise AI value gap

The signal

AI models now generate usable, well-structured output on demand, and they beat the average writer on grammar and polish. That is exactly why output quality stopped being the differentiator. Two operators running the same model get different value, and the difference is the system each one builds around it: how work is specified, what context the model sees, how output is judged, and who decides.

The enterprise numbers agree. PwC surveyed 4,454 chief executives in January and 56 percent reported zero cost or revenue improvement from AI in the prior year. The models perform in the demo. The value system around them is what most operations never built.

I run such a system daily across two publications and a product operation, all with AI in the loop. Here are the five practices that hold.

1. Specify the outcome, not the artifact

Most requests describe a deliverable. Write a post. Summarize this document. The model produces it fluently and still misses, because the deliverable was never the point.

State the effect the artifact must produce and the test it must pass. Before generating, I give the model a success definition: accurate, within limits, aligned with my written rules, useful to a named reader. Without a test, fluency becomes your quality bar, and fluency is the one thing every model already has.

When output misses, treat it as diagnostic evidence, not bad luck. A miss almost always reveals an unclear goal, a missing constraint, or a preference you never stated. Fix the specification and rerun once. My operating rules cap regeneration at one attempt per draft; the cap forces the diagnosis.

Do this: write the success test into the request itself. If you cannot state the test, you are not ready to generate.

2. Feed it what changes the answer

Context is leverage, but not by volume. A large undifferentiated dump buries the governing principle, because the model weights everything it is given. What matters is the small set of decisions, constraints, exclusions, and examples that materially change the answer.

In my operation this is a set of canon files: versioned documents holding settled decisions, voice rules, and standing constraints, read fresh by every automation run. The selection is the work. Deciding what the model must know is a judgment the model cannot make about itself.

Do this: write one standing context page: decisions already made, constraints that never move, one example of output you accepted. Reuse it in every session.

3. Separate the moves, keep the decision

One prompt that frames the problem, generates the answer, evaluates it, and decides has fused four moves, and you have surrendered three. Run them as steps: define the problem, generate options in the plural, critique them as a separate request, then choose yourself. The moment you accept the first option by default, the model is deciding, and it optimizes for plausibility, not your situation.

Do this: never end a working session on a generation step. End it on your decision.

4. Ask for the disagreement

Models default to agreeable elaboration. They refine the direction you are heading unless told to reopen it. Extract the disagreement deliberately: the strongest objection to the current direction, plus the materially different approach you have not considered.

The same mechanism works in reverse as a gap detector. Feed the model your own draft or plan and ask what is missing, unsupported, vague, or contradictory. This reliably beats another generated version, because it points the model at what it does best: reading with full attention and no ego.

Do this: before accepting any significant output, ask one question of the form "what would make this wrong?"

5. Labs receipt: make the checks deterministic

Every essay I publish must pass a written pre-flight checklist: spelling conventions, banned punctuation, structure, length bands, plus a set of judgment standards. The mechanical half has exactly one right answer per check, which makes model review the wrong tool for it.

That was the Labs hypothesis: the mechanical checks could be code. The build was preflight_lint.py, a small Python linter that runs before any draft reaches review. The finding: roughly 60 percent of the checklist left the model's hands entirely. Model review now covers only the judgment checks, where full-attention reading earns its cost.

Do this: split your quality checks into deterministic and judgment. Code the first set. Spend the model only where judgment is required.

The boundary underneath all five

Decision rights stay explicit: written rules define what the model may recommend, what needs my approval, and what it never decides. Nothing publishes or spends money without founder sign-off, and that rule lives where the models read it. Close the loop too: compare what the model recommended against what actually happened after you shipped, and record which assumption was wrong. The metric is value, not volume; more generated material can mean less value once review costs rise.

And the facts, evidence, and stakes that give the work authority come from the real operation. The model can examine and extend them. It never manufactures them. An invented receipt is worse than none, because every practice above assumes the inputs are real.

Takeaways

  • Write the success test into the request. No test, no generation.

  • Treat every miss as a specification defect. Fix the request, rerun once, stop there.

  • Keep one standing context page. Curate by what changes the answer.

  • Run framing, generation, critique, and decision as separate steps. Keep the last one.

  • Code the deterministic checks. Spend the model on judgment only.

INVENEW exists to help tech builders, operators, founders, and leaders turn AI from experiments into working systems.

Note: Third-party company and product names belong to their respective owners and are used for identification and illustrative reference only.

In partnership with

Make Your Clients Famous.

Book more client interviews without adding researchers or coordinators. PodPitch finds relevant shows, personalizes outreach from your team’s inbox, and follows up until interviews land. Give clients more visible momentum, stronger authority, and another reason to renew.