Most AI agent talk starts wrong. Teams compare models before they can name job. Smartest, cheapest, fastest, then bolt one onto vague assistant. Great demo. Nobody can say what it owns or who catches mistakes.
Buying model is not product judgment. Models do not ship products. Workflows do.
Start with job, not model
Start with work someone struggles to do. Name user, trigger, input, and outcome. “Help the team” is not job. “Turn support request into draft response for review” might be.
Polished answer can hide weak workflow. If nobody can explain what happens before prompt, after output, or when it is wrong, demo has not earned place in product.
What decision is blocked? What handoff is slow? What evidence must output include? What needs person approval?
Model potential is not accountability
Model gives potential. Product sets conditions: what model sees, what it may do, which step follows, who approves a consequence, and how failure gets handled. That is difference between plausible output and work people can own.
This matters when outcome has weight. Wrong release-note sentence is easy fix. Changed account access or sent customer message is different. “Model decided” is not answer. Ownership needs to be visible.
Think of model as contractor. It needs brief, permitted work, review path, and record. Give contractor keys to every room with no brief or record, then do not blame intelligence when outcome gets weird.
“Draft for approval” beats fake “handles support.” Do not promise more autonomy than workflow supports. People need to know when to trust output, correct it, or stop it.
Workflow reveals model value
Vague prompt mostly measures how model survives ambiguity. Defined workflow exposes useful differences. One may make messy customer language into clear drafts. Another may pull details from long documents. Cheaper option may handle routine classification while stronger model helps with hard exceptions.
Failures also become diagnosable. Was job ambiguous, source material missing, review step wrong, or model unsuited to task? Those are product questions. They lead to fixes.
Test against product criteria:
- Useful output for representative work, not only happy-path prompts?
- Response time fit moment where user needs help?
- Cost fit task volume and value?
- Express uncertainty when evidence is weak?
- Behave acceptably with allowed material and actions?
- Can team explain acceptance, rejection, or escalation?
No universal winner exists because no universal product exists. Pick after workflow is clear enough to test, not after leaderboard or sales call.
A stronger model earns place when it improves outcome users care about. It should create less review work, fewer dead ends, clearer next actions, or capability workflow could not support before.
AI cost control has same operational lesson: visibility and limits beat hope. If workflow cannot show where calls add value, it cannot tell whether premium model earns its cost.
Accountability is feature
Good agent products show source material when it matters. They mark drafts. They let people correct, approve, or stop work. When output causes trouble, team should know job, evidence, action, and intended decision-maker. If not, define workflow first.
Technical appendix: what a harness must do
Harness is machinery around model, not product strategy. It supplies current context, bounded tools, permission checks, action records, retry limits, enough state to inspect work, and tests based on representative tasks. Build enough for defined job. Improve it when real use exposes gap.

Comments load here
Discussion is powered by Giscus and loads from GitHub Discussions once you sign in.
giscus · awaiting mount