Back to Insights
Newsletter

Quanta Bits: AI Can Produce Work Quickly. Deciding What Is Good Enough to Use Is Harder.

AI makes generating answers and automating workflows increasingly easy, but production value depends on whether outputs are good enough to use. Teams need a clear definition of done, evidence-based tests, ownership, review costs, and controls proportional to what the system is allowed to do.

August 9, 2026

This issue covers two weeks of research and three patterns: company memory is becoming AI infrastructure, model selection is becoming one decision among many, and generating answers is getting much easier than judging whether those answers are good enough to use. All three move the difficult work into the controls around the model.

AI Can Produce Work Quickly. Deciding What Is Good Enough to Use Is Harder. - The main essay. After a year of building with AI and talking with business leaders and practitioners, one lesson stands out: AI can produce a great deal of work and automate processes quickly, at least in a pilot. The harder question is which workflows are safe and useful enough for production.

When a system can create thousands of answers, a vague definition of done becomes a much bigger problem. In one recent study, AI models generated complete bacteriophage genomes, but only laboratory testing established which designs produced viable viruses. In a more familiar business workflow, Built Technologies uses one AI step to extract information from finance documents, a second to compare the answers with the source, and people to handle uncertain cases. The important design choice is the evidence required and who owns the exceptions.

The definition of done often already exists in process guides, approval policies, required fields, brand standards, and experienced judgment. AI makes more of it testable. An amount matching a cited clause can become a check; non-standard payment terms can become a routing rule. Subjective work still needs sampling and measurement of downstream results, not just confirmation that an artifact passed its checks.

Controls should scale with authority. A brainstorming draft that a person will rewrite can tolerate occasional mistakes. A system that approves payments or sends advice directly to customers needs clear tests, evidence for each decision, a named approver, and a safe way to stop or reverse the action. More authority requires stronger proof.

The tests themselves also need scrutiny. OpenAI reviewed 731 tasks in SWE-Bench Pro and found roughly a third were flawed: some omitted requirements, some rejected valid solutions, and others let incomplete fixes pass. A clean score can still approve bad work when the test is poorly designed.

Review and correction belong in the return-on-investment calculation. In a small Mayo Clinic pilot, physicians using an AI scribe spent more time completing notes than those using a human scribe. The broader question is whether the value of work good enough to use exceeds the full cost of generation, review, correction, and failure.

Also in this issue:

  • Patterns & Signals - Shared agent memory is becoming company infrastructure, model choice is becoming a workload decision, durable systems are separating from replaceable features, and MCP is maturing into enterprise infrastructure.
  • Meanwhile... - Google DeepMind's WeatherNext model improves tropical-cyclone forecasts and gives forecasters roughly one additional day to prepare.
  • What I'm Consuming - AI wishing versus measurable results, an OpenAI engineer's Codex workflow, minimum viable governance, model collapse, and a chatbot powered by actual people.
  • On Digital Edge - A discussion about build versus buy, why the first version is usually the easy part, and how security, rollout, maintenance, ownership, and ROI determine whether custom software survives.
  • After Hours - The Manchurian Candidate, a fast-moving classic about overt brainwashing and the quieter manipulation of families, politics, and public opinion.

Before scaling an AI workflow, ask who owns the result and can stop it, how success will be measured, which existing steps should be eliminated instead of automated, and what evidence must accompany each output. If the automated process costs more or performs worse after checking, correction, and recovery are counted, it is not a saving.

Read the full newsletter on Beehiiv

Want More Like This?

Quanta Bits delivers curated automation insights to your inbox.