updated 2026.07.17 with Boris Cherny(Anthropic)'s "Steps of AI Adoption"

reading Yongho Ha, Hyangro, Toss team, and Anthropic side by side

New models arrive not monthly but weekly, and the experience of using AI shifts underneath us each time. I wanted to know how companies are actually doing AX — not the shallow kind, where AI Chat is adopted as one more SaaS subscription, but the deep kind, where an organization pushes AI transformation through its real work and comes out the other side with hard-won struggles, wins, and failures. I'm grateful that Yongho Ha, Hyangro (CTO of Inflearn), and the Toss team shared exactly those accounts.

To these I add a fourth source: Boris Cherny, who works on Claude at Anthropic, and his official blog post "Steps of AI Adoption." Where the first three are field reports ("here's what we — or the companies I watch — went through"), this one is a maturity ladder drawn from the tool-vendor's side by someone who talks to engineers at many companies every day. Because it's a different kind of thing, overlaying it pays off.

The problem Boris named in his tweet is exactly this piece's starting point: "I talk to engineers at other companies every day and hear the same thing: one person is 10x'ing their output with Claude but the rest of the org hasn't caught up."

Laying the four side by side, one thing jumped out. They stand in completely different places, and they hit almost exactly the same walls. How they get over those walls, though, depends on scale. So this piece doesn't stop at summarizing each source — it goes on to what only becomes visible when you overlay them.


0. Where each source is standing

If you want to transplant an insight, you first have to know what soil it grew in. That's why Hyangro spends the first third of his post on organizational context.

Hyangro (CTO, Inflearn) Yongho Ha (CEO DataOven / CDO Dable) Toss Team (TW Chapter) Boris Cherny (Anthropic)
Vantage point Runs a 56-person org Advises many companies Builds a knowledge system in a 4,000-person org The tool vendor, watching many companies
Scale 56 total N/A (many observed) ~4,000 community, 3 TWs N/A (Anthropic is at step 3→4)
Nature "Here's how we're doing it" "Here's which stage you're in" "Here's how we failed" "Here's how to climb this ladder"
Core asset Realism on infra, security, pricing A diagnostic frame (three debts, J-curve) Execution data and failure stories A capability ladder (0→4) + transition recipes
Bias to watch Small-org operator's pragmatism An advisor's cold diagnosis A practitioner's confession Vendor optimism (every rung maps to a Claude product)

Stitched into one sentence: Yongho Ha diagnoses the disease, Toss documents the course of treatment, Hyangro shows how the prescription changes in a small organization, and Boris draws the stage chart all the way to full recovery — just don't forget that the last chart was drawn by the company selling the medicine.


Part 1. Hyangro — "Let me start with our context"

An eight-month interim report from someone running a 56-person org (Inflearn / Rallit; 1.65M cumulative signups, 540K MAU) that began adopting AI seriously in November 2025. The first third is org specs — and that is the argument.

"When I read posts sharing AX insights, what's often missing is the context and environment the insight came out of. Every organization differs in size, budget, and infrastructure. Strip that premise away and even a good insight is hard to carry into your own situation."

Insight 1 — Efficiency and effectiveness are different claims

The first concession he makes, and the coldest coordinate among the four sources.

"I haven't seen a company above ₩10B in annual revenue — not GMV — grow 50%, 100%, or 200% through AX. What's been proven at this scale is 'you can produce your old output with fewer people.' I have not yet seen a case of 'you produce dramatically more.'"

His explanation for why it doesn't convert: AI only gets leverage where explicit knowledge already exists. If "do A → B → C and you get 100" is already codified, AI automates that path fast. But an organization that has never once produced 500 is not going to reach 500 with AI.

Case — the Tesla Model 3. Musk pushed for full automation early, production stalled, and in April 2018 he conceded: "Excessive automation at Tesla was a mistake. Humans are underrated." The principle he later articulated — "Question the requirements, delete the unnecessary parts, simplify, accelerate cycle time, and then automate" — says the same thing: automate an unvalidated process and you fail.

Insight 2 — A non-deterministic tool is useless to a passive person

"Tell an unmotivated teammate to 'do AX' and they'll use AI exactly as instructed, take no responsibility for the outcome. And that's when AI becomes the worst possible tool. It's non-deterministic — three tries, five tries, no guarantee. So it works for proactive people and does nothing at all for passive ones."

The warning that follows: the tighter the guidelines, the more people just execute them, and changing approach when output is wrong becomes nearly impossible. → "What's demanded of leaders most is motivating their people."

⚠️ This collides head-on with Toss's conclusion. Part 5 returns to it.

Insight 3 — The premise "anyone with money" is collapsing (the highlight)

This appears in no other source, and it's the part that landed hardest.

The original framing:

"If anyone with money can use the best model, what's the moat? This is not human vs. AI — it's organization vs. organization, both using the same tools."

But the premise is breaking:

"Top-tier models like Fable 5 are moving off flat-rate plans and onto metered pricing only. Until now, a small company got effectively unlimited tokens cheaply on a flat plan, while a large company was stuck on enterprise — ₩2–3M per person and still below a Max 5x. That asymmetry was actually a small company's advantage. Remove the top model from flat-rate and the line gets drawn between companies that can absorb metered cost and those that can't."

💰 A case that fell out of this: one company, forced onto enterprise, budgeted ₩1.5B a month and spent north of ₩4B — and is now re-examining AI usage across the board.

The question he leaves open: Inflearn runs $125 premium seats and $200 Max 20x. Fable-class models are out of reach there.

"If we're competing while using a lesser model than they are, what should we be doing?"

He does not answer this. He leaves it open. (My attempt is in Part 6.)

Insight 4 — The real intellectual asset is the decision process

"If you don't even know what your company's tacit knowledge is, start by recording every internal meeting. Documents preserve only outcomes and follow-ups. Your real intellectual asset is the decision process itself."

  • You need a framework for recording that process — like Google's ADRs (not a specific technology).
  • Without the reasoning, a slightly different situation produces a completely different decision. "And AI ends up with a completely different context too."
  • An alternative: embed a dedicated person inside the team who continuously extracts explicit knowledge AI can use.

🔧 The three biggest productivity lifts (measured at Inflearn): the Knowledge Base / a set of MCPs (Atlassian, Google Calendar, Gmail, BigQuery, Mixpanel, Datadog, GitHub) / FDE embedding — a backend engineer moving into marketing as an AX engineer.

Insight 5 — Consolidate onto one tool (diminishing returns don't bite here)

"We used to let everyone use whatever suited them best — the sum beat pooling know-how. But AI's possibility space is effectively unbounded, so diminishing marginal utility doesn't really apply."

→ Dozens or hundreds on one tool, concentrating plugins/skills/MCPs/prompt guides, beats individually-optimal tools. Fragment the tools and you fragment the know-how.

An often-missed point attached to this:

"The LLM is not holding your company's data. It translates your natural-language query into your document tool's search spec and calls it. So whether the thing you're looking for actually comes back depends more on your document tool's search quality than on the model."

Insight 6 — Infrastructure: not so AI can read it, but so you can go back

  • Git. "AI is non-deterministic — more iterations don't guarantee better results, and it often blows away work." If Git is too much, at minimum your document tool must version everything.
  • IaC & GitOps. Product teams change infra by PR while infra only reviews. "The psychological safety and sense of speed this gives you is a large, tangible difference." (Validates config, not actual cloud behavior — that needs integration tests.)
  • DB migrations (Flyway et al.). "Only versioned data lets AI understand intent and context — and lets you rewind and fix things when AI gets it wrong."

Insight 7 — Security: block it outright and it goes underground

"Say 'no' reflexively and people use AI in the shadows — a far bigger problem."

  • DevOps + AI + security in one org.
  • No personal accounts. Too many MCP-borne incidents — whitelist MCPs internally.
  • An AI proxy gateway is mandatory. "Traditional monitoring detects anomalies on request count. But an AI API can have identical counts and astronomically different cost — count-based monitoring will never catch that."
  • A real warehouse like BigQuery. A read-replica RDB needs individual GRANTs per account, unmanageable at hundreds of people. (On-demand bills by bytes scanned → consider flat-rate, partitioning, clustering.)

🚨 The scariest case — attack through the monitoring SDK. "There are cases where the attack comes in through the SDK that collects client-side errors. A case disclosed in June 2026 was exactly this — the attacker injected crafted data through an exposed error-collection key, and an AI coding agent, reading it through MCP, mistook it for a trustworthy instruction and executed it." (writeup) "Full-auto bug fixing with no human means one of two things: your service is too small to attack, or your security team is perfectly blocking everything." "Repository secrets and CI are almost never fully isolated between backend and frontend. Poison the frontend and the whole service is in scope."

🔗 This security lens is exactly why Boris's ladder (Part 4) has a separate "guardrail" column at every rung. Boris makes automatic code review, automatic security review, agent sandboxing, and worktree isolation the railings you must build to climb higher. What Hyangro felt as fear on the ground, Boris institutionalizes as a precondition of the ladder.


Part 2. Yongho Ha — "Which stage is your company in?"

Written from the vantage point of advising many companies, so the asset is patterns. Companies beginning AX go through remarkably similar stages and pains (denial → anger → bargaining → depression → acceptance).

The five-stage map

Stage What happens Symptom
1. Euphoria Company-wide rollout, outside trainers, an AX task force "This will transform us"
2. Stagnation Nobody uses it. Engineers a bit; non-engineers really don't Broken link: Input (hard to feed our context) / Output (paste results back into internal systems)
3. Excitement Pioneers (fighting security) connect internal systems. MCPs/Skills proliferate. An internal hackathon drives non-engineer usage up Token leaderboards and token maxxing
4. Doubt ← where most companies are now Underwhelming. Perceived speedup 10–20%. Nobody cares about the leaderboard anymore
5. The last hurdle Clear Pipeline Adaption, or cut headcount "Let's cut."

📐 This emotional/organizational ladder is a different axis from Boris's capability ladder (0→4) in Part 4. Overlaying them locates "where we are" in two dimensions. → Part 5-5 does this in full.

His balanced read on token maxxing:

"What we want is more output, but that's hard to measure. Token usage — an input — is 'the first intuitive, real-time proof of work' executives ever had, so managers love it. But by the later stages you must move to managing output. Token maxxing gets gamed, and companies clamp down on cost after being alarmed at spend without results."

🔗 Boris says the same thing (and he's the vendor!): "A usage dashboard is worth watching, but it measures activity, not return." The advisor (Yongho) and the supplier (Boris) name the leaderboard's limit identically. Boris's alternative metric is in Part 4.

What actually blows up in Stage 4:

  • AI-generated, under-reviewed code causes a production incident.
  • A decision on an unscrutinized AI report turns out wrong.
  • The alarmed company hard-rolls-back. → "When that happens, that company's AX is dead for six months."

This source's biggest asset — the three debts

The frame explaining the AI J-Curve Trap: Learning Curve + Verification Tax + Pipeline Adaption.

Debt Definition Why it grows in the AI era
Technical debt Output volume slows the next build AI code is locally optimal, globally ignorant. And plausible
Cognitive debt You can neither understand nor vouch for it Output arrives in a heap; understanding becomes a separate, deliberate act
Intent debt You can never again know why Working alone with agents, intent survives only in a prompt that evaporates

Technical debt — AI's debt is plausible:

"AI doesn't produce obvious bugs. Unit tests pass. It breaks when you wire the whole thing together in production. Code generated faster means debt accumulated faster." → Without a dedicated response, velocity gets worse within 5–19 months. (Simulation-derived; the trend matters, not the number.)

Cognitive debt — "cognitive surrender":

"It used to be that output arrived slowly, built by your own hands, so context happened naturally. Now output arrives all at once, and understanding has become a deliberate, separate act. And companies don't realize they have to budget time for it." → People read only the conclusion, give up on the rest, and hand it on. Who does the same. → Karpathy: you can outsource thinking but not understanding — except people surrender understanding too."A pipeline of one-click hand-offs: from your click to my click."

📌 A real case. A report came back: "there's no increase in news traffic." By AI's sense of an average it wasn't a spike — but in that company's domain, a 20–30% increase is an enormous inflow. Thinking got delegated, and the (wrong) result circulated internally.

Intent debt — worse than cognitive debt:

"Cognitive debt means it exists but wasn't understood. Intent debt means it never existed at all." "The person who knew that left." / "The reason was in a chat thread and I can't find it." Teams used to back up context in each other's heads; the further AX goes, the more one person works alone with agents and that trail disappears.

📌 Case — the companies that rolled back. Firms that cut headcount preemptively rehired the same people at higher salaries (Google, Salesforce, Duolingo, Klarna, CNET). → "So far, a human head is still the best storage device we have for tacit knowledge."

The remedy — move human work from production to verification

The key is not verifying everything: don't verify code and intermediate states — verify the result rigorously.

The duck test: "Wherever it came from, however wildly the AI thrashed, if it passes hundreds of verification layers a human carefully defined, you can trust it. And the tacit knowledge you inject while building those layers pays down intent debt."

Three kinds of layer: (1) Binary checks — tests; (2) Quantitative metrics — throughput, latency; (3) Qualitative rubrics — "extensible without over-abstraction?", "too many colors?" via LLM as a judge (1–5). And verification is needed at run-time, not just build-time — if the product is an agent, it occasionally goes off the rails.

🔗 Boris's entire ladder runs on this "verification" engine. The condition for step 1→2 is "a self-verification loop you trust (tests + build + lint + e2e)," and the bottleneck for step 2→3 is "trust in the loop." Yongho's 'verification layer' is Boris's 'trust in the loop.' Only the name differs.

🧠 The moment perception shifted — the Claude Code source leak. The source of the tool that changed how the world writes software leaked, and the quality was lower than expected. People realized: "We needed A-grade code all this time to fit human cognitive space. If AI can freely handle un-organized code and just delivers the result, isn't that enough?""If it passes verification and the result is guaranteed, the process doesn't have to be optimal."

💡 A usable tip — always verify in a separate agent. "An LLM defends what it just said. Ask the same session to 'also verify' and it catches only a fraction." A prompt that works: "Spawn N sub-agents that critique the existing work in parallel from different critical perspectives. If a critique holds up, re-verify in the main session and accept if valid. Repeat twice."

Extracting tacit knowledge — reverse the roles

Tacit knowledge is "what I know but don't know that I know." So let AI do the asking.

Matt Pocock's "grill-me" (github.com/mattpocock/skills): "Interrogate every aspect of this plan relentlessly… ask one question at a time… if the codebase can answer it, check the codebase first instead of asking me." "grill-with-docs" — same process, but emits Markdown the next session reuses.

"Conceptualize yourself as the verifier and advisor, and set AI up as the aggressive questioner. The questioner is not you — it's the AI."

The conditions for an AI-native company — and the real bottlenecks

Conditions: Queryable → Closed loop → Self-improving.

🔗 These three are exactly Boris's definition of step 4 "AI-native" ("the loop is fully closed"). The advisor and the vendor independently drew the same destination.

But once you have that, is velocity unbounded? No. Three bottlenecks remain:

  1. Building an SSOT is extremely hard. "Most haven't even digitized, and drift creeps in over time."
  2. Burnout arrives fast. "However good the layers, there's always a human-in-the-loop element, and that human becomes the bottleneck.In the AX era, managing mental energy paradoxically becomes the most important thing."
  3. Taste won't converge. ← the real problem. "Human taste-convergence runs on slow meetings, so it, not AI's generation speed, becomes the bottleneck."

⚠️ None of these three appear on Boris's ladder. Boris charts how to close the loop through step 4, but not the human bottlenecks that remain after it's closed (burnout, taste convergence). The blind spot of vendor optimism is exactly here. → Part 8.

So what's left for humans

In an 8-hour day: before AI = 6h doing + 2h judging → after = 2h instructing + 6h judging. The role model is your CEO — worse than you at everything, yet directs the work and runs the company. AI will soon be better than you at engineering, marketing, design — and you must become the one who directs it.

🔗 This lines up with the human roles Boris pins on each rung: step 1 "you + an agent (a pair)" → step 2 "Orchestrator" → step 3 "Manager of managers" → step 4 "VP steering by intent." Yongho's "become the CEO" is Boris's "VP steering by intent."

Three capabilities of "someone who owns a job start to finish": (1) decompose problems, (2) detect failure fast, (3) find the structure that makes work work. → In short: "the ability to find answers in ambiguous situations."

Why expertise still matters — the Gell-Mann amnesia effect: the medium changed from newspapers to LLMs; the bias is identical. Some of what looked plausible from AI looked that way because we weren't experts.To filter the plausible fake, you need expertise. The domain expert builds the most valid verification layer.

The conclusion: "From master of a skill to owner of an operation."


Part 3. The Toss team — "Here's how we failed"

A six-part series. 4,000 people; three TWs. The value is the trajectory of failure. Each stage exists because the previous one failed.

① One person writes more        → failed: can't keep pace with a changing product
② Build a culture (workshops)   → failed: the first contribution happens, the second doesn't
③ Ship a tool (a Skill)         → failed: nobody used it — you had to invoke it deliberately
④ Embed it in the workflow      → worked. And now too much has piled up
⑤ Set standards and governance  → they are here now (Knowledge Committee)
⑥ Knowledge circulates itself   → the goal (create · verify · refresh · retire)

Most organizations stop somewhere in ①–③ and conclude "our team just lacks the will." The value of this series is that it proves, with failure data, this was never a will problem — it was a structure problem.

Insight 1 — Code is not the SSoT

"Code preserves only the result. It records what it does, not why. A true SSoT is only complete when the code and the context around it are both present."

Their definition of knowledge: "verified information that helps someone make a better decision." → If it isn't verified, it isn't knowledge.

Insight 2 — Documentation fails on structure, not on will

"Documentation is work for 'later,' not 'now,' so it's always deprioritized." "Documentation always begins with individual willpower. The writing is quick — but the decision to write is expensive, so it slips when you're busy."

The cost of missing docs: you spend more time hunting history than developing; unknown unknowns last far longer; a psychological hurdle ("everyone probably knows this / maybe I'm the only one who doesn't").

Insight 3 — Lower the burden and behavior changes (measurably)

The problem: "Asking a question means publicly admitting you don't know something."

Intervention Result
A "consult the bot" week Questions 2.5/day → 11/day (4x+); 40 more people joined
A bot broadcasting one small fact daily "Being told to write a finished doc is daunting. Adding a line to shared knowledge is far easier."

Insight 4 — A question the bot couldn't answer = a doc is missing ⭐

Across all four sources, the highest-leverage idea per unit of cost.

Someone asks the bot → the bot can't answer → "there's no doc here"
                                            → AI gathers sources and drafts one
                                            → a human just checks the evidence and approves

Documentation's hardest problem is "I don't know what to write." This loop harvests that from users' actual questions. Toss Commerce went further: every night, AI drafts from two signals — deploy/policy announcements, and questions the bot failed to answer. A human only checks and approves.

Insight 5 — A checklist can be poison to an AI ⭐

  • Attempt: review comments → checklist → AI checks item by item
  • Result: "It missed the problems that mattered and forced out comments that didn't need to exist."
  • Cause: "The conditions for good writing are fixed. Bad writing is broken differently every time."
  • Fix: give principles + (incorrect ↔ correct) pairs and let AI judge.

Generalized: where the right answer converges, checklists work. Where the wrong answer diverges, checklists degrade performance — supply the reasoning and leave room for judgment.

Insight 6 — The moment a tool needs deliberate invocation, adoption dies ⭐

Toss built a good Skill, shipped it, and nobody used it. Three reasons: hard to install; you had to consciously remember mid-work; a human had to find and hand over the source material.

The fix: don't move people to the tool. Plant the tool where people already are.

Task Where Why
Writing The messenger Tacit knowledge surfaces in conversation
Editing GitHub PRs The moment a review begins is already defined

🔗 Boris's step 3→4 is the endgame of this principle. Higher up, "most agents are kicked off by Claude" — the human doesn't press a trigger; Claude monitors a channel or data source and starts work itself (Tag, routines, /loop). Toss found an existing trigger (the PR); Boris hands the trigger itself to AI.

Insight 7 — Knowledge becomes debt the moment you accumulate it ⭐⭐

The final twist. Toss succeeded at documentation, and because it succeeded, met a new problem.

  • "Too much accumulated." / "The bot answered, calmly, on a stale policy — and was wrong." / "The bot presented a finished experiment as current policy." / "When a policy owner changed, nobody could explain why the criteria were set that way, and the overhaul stalled."

"Tools gather info faster. But 'can I trust this,' 'who owns it,' 'is this current or a finished experiment' — a tool can't decide that. Automation didn't remove the problem; it made the absence of standards impossible to ignore."

Knowledge has a four-stage lifecycle; most automate only the first:

Stage Difficulty Most orgs
Create Easy
Verify Medium ⚠️ leans on human approval
Refresh Hard ❌ usually absent
Retire Hardest (needs owner + standard) ❌ essentially absent

Toss's answer: governance. A Knowledge Committee — unlike a guild, "a defined membership holds decision authority and its decisions carry real force." Two layers: TW chapter sets company-wide standards; each domain operates them locally. Four standards: leave what you know to the org (repeated questions / key decisions / new-joiner essentials — unconditionally); make it findable (human + AI); make it usable in the workflow; keep it current with an owner and a review cadence.

"Keeping knowledge trustworthy and usable matters more than accumulating it."

Insight 8 — Split docs for humans and docs for AI

"Docs AI reads often need detail humans don't."

Commerce split into central docs (human-friendly, TW-managed) and per-team repos (AI-friendly, auto-accumulated). Agonizing over "should this AI-only context go in the shared doc?" was itself the hurdle. "This structure exists precisely because documents are no longer read only by humans."


Part 4. Boris Cherny (Anthropic) — "Here's how to climb this ladder"

Where the first three ask "what's the problem / how did we fail," Boris's "Steps of AI Adoption" (2026.07.16) charts "which rung you're on and what it takes to reach the next." It's a map the tool vendor drew from watching many companies, so it reads differently from a field report.

The one line that runs through it (the core)

"Tokens aren't enough to move you forward: to get to the next step, you need to find and break down the next set of bottlenecks, and build up the next set of guardrails."

It matters that the vendor said this. The side selling tokens (= revenue) insists tokens alone won't do it. And those "bottlenecks + guardrails" are precisely what the first three sources call context and verification. (→ synthesized in Part 6.)

The capability ladder (0 → 4)

Step Human role Agents What it looks like Bottleneck Guardrail
0 Gated 0 Only older/lighter models approved; no MCP governance; access gated or process-heavy; outputs exist only locally Legacy security/approval; fixation on cost-per-token, not outcomes; no true technical voice in decisions SSO/SCIM, org budget caps, deploy inside existing approvals/IAM
1 Assisted
you + an agent (a pair)
Pair programmer ~1 One engineer, one agent, mostly supervised. Review almost every change before merge. An afternoon's task finishes between meetings Your attention. You don't trust the output so you read everything and never look away. Work is synchronous Per-seat caps, central model/policy, OpenTelemetry into SIEM, Plan mode
2 Parallel
Orchestrator
Orchestrator ~10 One engineer orchestrates 5–10 agents on separate worktrees. Claude self-checks (test·build·lint·security) before you see it. Auto mode always on; automated code/security review on by default. You review final diffs, not keystrokes. Backlog starts shrinking Reviewing output. You check six streams instead of writing. Prompting/steering while juggling Auto code-quality (lint·test·typecheck), Claude-powered e2e verification, same bar for human and agent code, pre-approve safe bash/MCP
3 Supervised autonomy
manager of managers (org tree)
Org-tree manager ~100 Claude writes all or nearly all code. "Did you read the code?" becomes "what context was the model missing, and how do we fix it next time?" Maintenance runs continuously in the background Trust in the loop and team decision throughput. The tree is too deep to babysit. The trap: scaling agent count before the loop earns trust. Token efficiency Auto code/security review, agent sandboxing, CLAUDE.md + Skills to encode standards, tune the Auto-mode classifier
4 AI-native
VP steering by intent
Steering by intent ~1,000+ The loop is fully closed and most agents are kicked off by Claude. Hundreds–thousands run; you steer by intent and monitor by exception. A quarter-long migration becomes a workflow you kick off and check on Identifying and automating work at scale, and enforcing the right guardrails per work type Cost controls for automation, model selection, Agent SDK to build/schedule agents

Step-transition recipes (a practical checklist)

  • 0 → 1: Executive/buyer alignment and blocker escalation; a framework to launch Claude securely
  • 1 → 2: run more than one agent at once; a self-verification loop you trust (tests + build + lint + e2e in a real dev env); auto mode so permission prompts don't block; automate code review
  • 2 → 3: let Claude pull in context itself (read code, wikis, discussions); review speed and agency (agents touch other teams' code); break work into loops/routines; let Claude kick off Claude
  • 3 → 4: scaled automation of domain-specific use cases (code migration, fuzzing, feature-building, feedback remediation)

How to measure ROI (a vendor's answer to Yongho's problem)

"Usage is worth watching, but it measures activity, not return. A better question: would you have spent engineering effort on this anyway? If yes, how much, and what would it have cost in manual eng-hours? That's your return."

  • I.e., convert not to tokens/usage but to "manual engineering hours displaced."
  • His self-report: "Anthropic is on step 3 and pushing toward 4. Personally, I just hit level 4."The gap between the org (3) and the individual (4) is the tweet's whole thesis ("one person 10x's, the org hasn't caught up").

A caveat while reading this (bias)

Every "Products that help" cell is a Claude product (Auto mode, Agent view, Code Review, Security Review, Tag, Agent SDK…). The ladder is an insight, but its rungs are all mapped to the vendor's own products — read accordingly. And the human bottlenecks that remain after the loop closes (Yongho's burnout, taste convergence) aren't on this ladder. That's the blind spot.


Part 5. What you see only when you overlay them

5-1. The three field reports hit the same wall

(Boris sits out this table — he charted a ladder, not a wall. Part 5-5 overlays him directly.)

The wall Yongho Ha Hyangro Toss
Context disappears Intent debt "no reasoning → AI gets a completely different context" "Code preserves only the result"
You can't trust the output Cognitive debt "non-deterministic; more tries don't guarantee better" "The bot answered calmly on a stale policy — and was wrong"
Verification is the answer The duck test Tests, AI review, integration tests in infra "Knowledge is verified information"
The SSOT won't hold "Building an SSOT is extremely hard" "Record meetings. Write ADRs." todoc

So whether it's 56 people or 4,000, the diagnosis converges.

For AI to do well it needs the organization's context. That context lives in people's heads and in prompts that evaporate. And extracting it will never be sustained by individual willpower.

5-2. And here they collide — motivation, or structure?

Hyangro: "AI works for proactive people and does nothing for passive ones. What's demanded of leaders most is motivating their people." Toss: "Documentation fails not from lack of will, but from a structure that depends on individual willpower."

Both came out of real practice. The reconciliation: they operate at different layers.

  • Structure creates the surface on which motivation acts. Anything that only happens if someone wills it must be pushed into structure. → Toss is right.
  • But a layer remains that structure can't reach. Trying five, ten more approaches when AI fails you can't be mandated. → Hyangro is right.

So the question isn't "motivation or structure" — it's "what to push into structure, and what to leave to motivation." Toss's data proves it: the first workshop doc happens; the second and third don't.

The first one happens through motivation. The tenth only happens through structure.

Saying "let's get better at documentation" bets the tenth on motivation, and that fails without exception. A leader's job is to make the tenth structural, and protect the motivation that produces the first. The same failure recurs in Yongho's "the human-in-the-loop becomes the bottleneck" and Toss's "we automated creation and left refresh/retirement to will."

5-3. Governance is a precondition for Queryable

To be Queryable        → knowledge must exist
For knowledge to exist → it must be trustworthy    ("if it isn't verified, it isn't knowledge")
To be trustworthy      → someone must own it

This is the true identity of Yongho's "SSOT is extremely hard" bottleneck. An SSOT is not a technology problem — it's an accountability problem. Toss built the tool (todoc) and still had to build a structure of accountability (the Knowledge Committee).

5-4. DX before AX — but purposeful DX

Yongho: "Every company's data comes in two kinds: it doesn't exist, or it's unusable." DX's central flaw was always DX without a purpose.

Hyangro's infra list is exactly "DX with AX as its purpose" — adopted not "so AI can read it" but "so we can go back when something goes wrong." For an org using a non-deterministic tool, the infra requirement isn't fast — it's reversible.

5-5. Two ladders — capability (Boris) × emotion (Yongho) ⭐

The highest-value part of adding Boris. Both drew a "stage chart," but the ruler is completely different.

  • Yongho's 5 stages = emotional/organizational temperature (euphoria → stagnation → excitement → doubt → hurdle). "How we feel right now."
  • Boris's 5 steps = capability/automation level (Gated → Assisted → Parallel → Supervised → AI-native). "What we can do right now."

They're orthogonal. Overlay them and you get a 2-D coordinate.

            Low capability ───────────────► High capability
Good mood     [Euphoria]                      [The dream of full AX]
   │       just adopted, excited              loop closed, actually running
   │
   │       [Stagnation]      ┌──── the J-curve pit ────┐
Bad mood    nobody uses  →   │  Boris step 2→3 zone    │
            [Doubt]          │  = Yongho stage 4       │
                             │  bottleneck: trust      │
                             └─────────────────────────┘

The key finding: the location of Yongho's J-curve pit (stage 4 "doubt") coincides exactly with Boris's step 2→3 bottleneck ("trust in the loop").

  • Boris names it coldly: "the trap is scaling agent count before the loop has earned widespread trust" — which is precisely the mechanism of Yongho's stage-4 incident (unreviewed AI code causing a production failure).
  • The moment the capability to run 10, 100 agents (Boris) outruns the verification to trust the result (Yongho), the company falls into the pit.

Merge the two ladders and you get:

Climb Boris's ladder, but pre-pay Yongho's "Verification Tax" at every rung. Build that rung's railing (verification, governance) before raising capability (agent count). Climb without railings and you fall as far as you've climbed (the J-curve).

This is no coincidence. It's exactly why Boris put a separate "guardrail" column on each rung and why he wrote that "tokens aren't enough to move you forward." The vendor and the advisor drew the same picture without knowing it.

5-6. Even the vendor says "the model and the tokens aren't the moat" ⭐

The strongest cross-check. The person at the company that sells tokens writes:

"It's not about a single feature, but using the right features with the right guardrails. Tokens aren't enough."

This is covering fire from behind enemy lines for Hyangro's Part-6 question ("what do we win with, on a lesser model?"). If even the model supplier says leveling up comes from removing bottlenecks and building guardrails (= context and verification) rather than the sheer volume of model/tokens, then Part 6's conclusion — that model tier is not the moat — gets much sturdier.


Part 6. Attempting the question Hyangro left open

"If we're competing while using a lesser model than they are, what should we be doing?"

Having read all four, the question resolves if you rephrase it.

Model tier is a constant we don't control. So what are the variables we do control?

Two. The context that goes in, and the verification that filters what comes out. Nearly everything all four sources describe reduces to these:

  • Hyangro's ADRs, meeting recordings, KB, MCPs → context; his tests, AI review, integration tests → verification
  • Yongho's three debts → the absence of context (intent debt) and verification (cognitive debt); the duck test → verification
  • Toss's todoc → context; Knowledge Committee → the trustworthiness of context = verification
  • Boris's "bottleneck (= context supply)" and "guardrail (= verification)"exactly these two variables. The vendor confirms it directly.

Why this is the answer. A model's capability determines how well AI reasons. But results depend on how well AI produces the right thing on your context. Without your explicit knowledge in order, a better model just generates a more sophisticated wrong answer. (Yongho's "no increase in news traffic" is exactly that.) With context and verification, a lesser model's output gets filtered and corrected into something trustworthy. Layered on: "If it passes verification and the result is guaranteed, the process doesn't have to be optimal."

A better model produces a more elegant process. But what guarantees the result is the verification layer — and good verification layers come from domain expertise, not money.

So, to put it together:

  1. Concede the model gap up front. You can't build strategy on an uncontrollable variable. (And even Boris says tokens won't carry you to the next rung.)
  2. What closes the gap is context — and it's a function of years of accumulation, so it can't be bought back overnight.
  3. What makes the gap irrelevant is verification (= Boris's guardrails). Without it, you can't trust even the best model's output.
  4. And both come from people. Back to hiring, training, and motivating — the conclusion Hyangro had already reached, except now you know why.

Anything you can buy with money doesn't differentiate you. Only what money can't buy becomes a moat.

Models you buy. Context, verification, and people you don't.


Part 7. What to take, by scale

Don't transplant these wholesale — exactly as Hyangro warns.

Small (~50 people) — a realistic target is Boris steps 2–3

Take this Source
The "bot couldn't answer = doc is missing" loop Toss. A nearly free signal. Top priority
Meeting recordings + ADRs Hyangro, Yongho. Start today with no tooling
Consolidate onto one AI tool Hyangro. Fewer people → bigger payoff from concentration
Verify in a separate agent Yongho. Fits in one prompt
Principles + examples, not checklists Toss. Just edit your review bot's prompt
Build the step 1→2 railing first: a self-verification loop + auto code review Boris. Before raising agent count
Do NOT take this Why
A Knowledge Committee A 4,000-person solution. At 50 the committee is the overhead. Use "one page of standards + a named owner."
A dedicated TW org No headcount. Use FDE embedding instead
Building your own doc platform (todoc) Half a year + a product team. Don't build — buy good search and consolidate
Step 3→4 scaled automation Even Anthropic is only at step 3. A 50-person team chasing a "fully closed loop" is J-curve suicide

Mid-to-large (300+)

  • Governance before tooling. Toss discovered "we have no standards" after building the tool; at scale that's far more expensive.
  • Splitting human/AI docs becomes mandatory.
  • Going metered without an AI proxy gateway will cause an incident. The ₩1.5B → ₩4B case.
  • The token leaderboard is a Stage 3 (Yongho) / step 1–2 (Boris) instrument. If you're higher and still watching an input metric, you missed the move to output/return. Pair Boris's return metric (convert to manual eng-hours).
  • Use Boris's transition recipes as a roadmap, but build each rung's railing first — guard org-wide against scaling agent count without "trust in the loop."

Part 8. What none of them could answer

The questions that remain. I think this is the actual frontier.

  1. Where is the crossover from efficiency to effectiveness? Hyangro: "haven't seen it." Yongho names "Pipeline Adaption." Boris gives a way to measure efficiency (manual eng-hours) but not when efficiency turns into revenue. Yongho's diagnosis stings: "feature velocity was never the constraint on revenue — direction, marketing, sales matter more, and those are still trapped in old workflows nobody dares reform."

  2. Who owns retiring knowledge? Even Toss is "still designing it." Create/verify/refresh can be tooled; retiring is a judgment, and judgment is accountability. (Boris's ladder has no "retire" rung either.)

  3. Who protects the human at the end of the verification chain? Yongho's second bottleneck: that human becomes the choke point. → The paradox that mental/energy management becomes more important in the AX era. Boris's ladder doesn't touch this human bottleneck at all (the blind spot of vendor optimism).

  4. How do you speed up taste convergence? The bottleneck Yongho calls "the real problem." (Hyangro's 1–4 person squad experiment reads as one response — reduce the number of people whose taste must converge.)

  5. What happens when top models leave flat-rate plans? The question Hyangro raises and the others don't touch. Ironically, the party driving that shift is Boris's employer.

  6. How does the org catch up to the individual? — Boris's new question. His starting point ("one person 10x's, the org hasn't caught up") and his self-report (individual level 4, org step 3) remain unresolved. Moving individual maturity into organizational maturity — this was the whole subject of Toss's six parts and of Yongho's "motivation" debate. The four sources are asking the same question in different languages.


Sources

The Toss series and Yongho Ha's talk are in Korean; quotations are my translations. Boris Cherny's post is in English; quotations are verbatim.