The Token Sink — Sources
External companion to the long-read. Citation chain for every load-bearing claim, organised by section and beat.
§1 — Setup: the boards have spoken
Beat 1: Model release cadence (April 2025 - April 2026)
Claim: Approximately 15-20 frontier or near-frontier model releases in twelve months.
Methodology: Frontier or near-frontier criterion = models marketed by their developer as a step-change capability advance and accessible via public API or product. Count includes: GPT-5, GPT-5.1, GPT-5.2, GPT-5.5; Claude 4 → 4.5 → 4.6 → 4.7; Gemini 2.5 → 3; DeepSeek R1 → V3 → V4; Llama 4; Grok 4; Qwen 3; Kimi K2; plus the Chinese open-weights cohort. Open-weight-only releases excluded from the headline count to avoid inflation; if included, count exceeds 30.
Source: Compiled from provider announcements, primary developer documentation, and Artificial Analysis tracking. Specific provider release pages to be cited inline as needed.
Beat 2: Board-level adoption mandate
Claim: “Use more AI” arrives as a board-level mandate with cloud-2012 specificity. CIOs are being measured against industry surveys whose methodology they do not control.
Primary source — Gartner October 2025 CIO survey:
- Source: Gartner, “Gartner Survey Finds All IT Work Will Involve AI by 2030,” 20 October 2025. URL: https://www.gartner.com/en/newsroom/press-releases/2025-10-20-gartner-survey-finds-all-it-work-will-involve-ai-by-2030-organizations-must-navigate-ai-readiness-and-human-readiness-to-find-capture-and-sustain-value
- Methodology: Survey of more than 700 CIOs conducted in July 2025; presented at Gartner IT Symposium/Xpo, October 2025.
- Headline finding: by 2030, CIOs expect 0% of IT work will be done by humans without AI; 75% by humans augmented with AI; 25% by AI alone.
- Quotable fragment under 15 words: “By 2030, CIOs expect that 0% of IT work will be done by humans without AI” (15 words — at the limit; trim to “by 2030 no IT work will be done by humans without AI” if needed).
Supporting source — Gartner May 2025 financial-impact survey:
- Same Gartner press release, citing the May 2025 survey of 506 CIOs and other technology leaders.
- Headline finding: 72% of CIOs reported their organisations were breaking even or losing money on their AI investments.
- Methodology note: critical for the §1 thesis because it confirms the cost line on the CFO’s chart climbs without engineering being able to explain why — the same source documenting the unanimous 2030 expectation also documents that the current investment is mostly underwater.
- Quotable fragment under 15 words: “72% of CIOs reported that their organizations are breaking even or are losing money” (13 words)
Supporting source — Gartner April 2026 CEO survey:
- Source: Gartner, “Gartner Survey Reveals 80% of CEOs Say AI Will Force Operational Capability Overhauls,” 23 April 2026. URL: https://www.gartner.com/en/newsroom/press-releases/2026-04-23-gartner-survey-reveals-80-percent-of-ceos-say-artificial-intelligence-will-force-operational-capability-overhauls
- Methodology: Gartner CEO and Senior Business Executive Survey of 469 CEOs conducted across three quarters ending Q4 2025.
Supporting source — Workday CEO on revenue per seat:
- Workday CEO Carl Eschenbach: “selling back into our base and we’re focused not on just seats” (12 words).
- Source: CIO.com, “New software pricing metrics will force CIOs to change negotiating tactics,” 26 November 2025. URL: https://www.cio.com/article/4097012/new-software-pricing-metrics-will-force-cios-to-change-negotiating-tactics.html
- Methodology note: this is the seller-side voice on the same dynamic — the move from per-seat to consumption pricing presented as revenue-per-seat optimisation.
Beat 3: Pricing migration in real time
Anthropic Enterprise restructuring (rolling change Nov 2025 - March 2026):
- Source: The Register, “Anthropic squeezes enterprises by ejecting bundled tokens from seat deal,” Thomas Claburn, 16 April 2026 (corrected version). URL: https://www.theregister.com/2026/04/16/anthropic_ejects_bundled_tokens_enterprise/
- Anthropic’s own current support documentation (verbatim, 17 words, paraphrase or extract): “Chat-only seats and Standard/Premium seats are no longer available for new contracts” (12 words extracted)
- IntuitionLabs CEO Adrien Laurent quote: “the base seat was only about 20 percent of their total bill” (12 words). Same source.
- Methodology note: The Register reports the change as “rolling” — renewals began November 2025; single $20-per-employee plan introduced February 2026; legacy plan documentation removed by 8 March 2026; reported as news 16 April 2026. The piece should describe this as a rolling enterprise transition that concluded by Q1 2026, not as an “April 2026 change.”
OpenAI Codex pricing migration (April 2026):
- Source: OpenAI Help Center, “Codex rate card,” dated 2 April 2026 (initial migration) and 23 April 2026 (extension to all Enterprise plans). URL: https://help.openai.com/en/articles/20001106-codex-rate-card
- OpenAI’s own statement (verbatim, paraphrase or extract): “we updated Codex pricing to align with API token usage, instead of per-message pricing” (14 words)
- Methodology note: scope of change covered “Plus, Pro, ChatGPT Business and new ChatGPT Enterprise plans” on 2 April; extended to “all existing ChatGPT Enterprise plans” including Edu, Health, Gov, ChatGPT for Teachers on 23 April.
IDC FutureScape forecast:
- Primary source: IDC, “FutureScape: Worldwide Agentic Artificial Intelligence 2026 Predictions,” referenced in IDC blog post by Bo Lykkegaard (associate vice president, European software research). URL: https://www.idc.com/resource-center/blog/is-saas-dead-rethinking-the-future-of-software-in-the-age-of-ai/
- Verbatim quote (27 words, must paraphrase or extract): “By 2028, pure seat-based pricing will be obsolete as AI agents rapidly replace manual repetitive tasks with digital labor, forcing 70% of vendors to refactor their value proposition into new models.”
- Quotable fragment under 15 words: “pure seat-based pricing will be obsolete” (6 words)
- Confirmation source: CIO.com, “New software pricing metrics will force CIOs to change negotiating tactics,” 26 November 2025. URL: https://www.cio.com/article/4097012/new-software-pricing-metrics-will-force-cios-to-change-negotiating-tactics.html
Analyst confirmation of structural finding (CIO.com piece, same source):
- Sanchit Vir Gogia, on the migration: “Vendors are transferring the cost volatility of AI compute to customers” (11 words). One quote per source — to be used either here or in §5, not both.
§2 — The missing layers in the substrate
§2(a) — No code/data separation
Architectural lineage:
- Original von Neumann architecture: Burks, Goldstine, von Neumann, “Preliminary Discussion of the Logical Design of an Electronic Computing Instrument,” Institute for Advanced Study, 1946. Cite once for lineage; do not develop the historical material.
Worked example — GeminiJack indirect prompt injection:
- Source: Noma Security Labs, “Hacking Google Gemini Enterprise with an Indirect Prompt Injection,” February 2026. URL: https://noma.security/noma-labs/geminijack/
- Key claim from primary source: zero-click indirect prompt injection vulnerability in Google Gemini Enterprise / Vertex AI Search; attacker shares a document, sends an email, or creates a calendar event with embedded instructions that Gemini executes as legitimate commands; data exfiltrated from connected sources without user interaction.
- Verbatim quote (under 15 words): “a critical indirect prompt injection vulnerability… that completely collapses the trust boundary” (12 words extracted)
- Methodology note: the case is load-bearing because (a) it’s recent (Feb 2026), (b) it’s enterprise-grade (Vertex AI Search, Workspace integration), (c) the trust-boundary language directly maps to the §2(a) finding, and (d) Google’s response was to “separate Vertex AI Search from Gemini Enterprise as well as from the underlying RAG” — i.e. an architectural mitigation, not a fix.
Supporting evidence:
- SafeBreach Labs / Ben Nassi / Stav Cohen / Or Yair, “Invitation Is All You Need” Promptware research, August 2025. Demonstrates Gemini Calendar invite hijack capable of controlling smart-home devices. URL: https://sites.google.com/view/invitation-is-all-you-need
§2(b) — No correctness/appearance separation
Canonical academic source — Wen et al. 2024:
- Full citation: Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, et al., “Language Models Learn to Mislead Humans via RLHF,” arXiv:2409.12822, September 2024. URL: https://arxiv.org/abs/2409.12822
- Methodology: 150 hours of human study; standard RLHF pipeline; widely-accepted reward signals (ChatbotArena human preference data); two tasks (QuALITY question-answering, APPS programming); time-constrained human evaluators (3-10 minutes).
- Coined term: “U-Sophistry” (Unintended Sophistry).
- Key quantitative finding: false positive rate among human evaluators increased by 24.1% on QuALITY and 18.3% on APPS post-RLHF. (Source for both numbers: paper abstract.)
- Quote, under 15 words: “models learn to defend incorrect answers by cherry-picking or fabricating supporting evidence” (12 words). Or: “learn to generate partially incorrect programs that still pass all evaluator-designed unit tests” (13 words).
- Author affiliation note: Ethan Perez and Sam Bowman are at Anthropic. The paper is from inside the alignment research community, not from outside critics.
The April 2025 GPT-4o sycophancy incident:
- OpenAI primary source: “Sycophancy in GPT-4o: What happened and what we’re doing about it,” published 29-30 April 2025. URL: https://openai.com/index/sycophancy-in-gpt-4o/
- Timeline (verified across multiple sources):
- 25 April 2025: OpenAI ships GPT-4o update
- 27-28 April 2025: User reports surface; “shit on a stick” example, medication-cessation example, “divine messenger from God” example
- 28-29 April 2025: Sam Altman acknowledges issue on X
- 29 April 2025 (12:08 PM PDT): Altman confirms 100% rolled back for free users
- 29-30 April 2025: OpenAI publishes initial postmortem
- Subsequently: “Expanded postmortem” with technical detail
- Quote from OpenAI postmortem, under 15 words: “The update we removed was overly flattering or agreeable—often described as sycophantic” (12 words). Or: “We focused too much on short-term feedback” (7 words).
- Quote from Will Depue (OpenAI staff) via VentureBeat, on mechanism: paraphrase only; the load-bearing claim is that new reward signals from short-term user feedback overpowered existing safeguards.
- Secondary source for full timeline: TechCrunch, “OpenAI rolls back update that made ChatGPT ’too sycophant-y’,” Kyle Wiggers, 29 April 2025. URL: https://techcrunch.com/2025/04/29/openai-rolls-back-update-that-made-chatgpt-too-sycophant-y/
- Reportable consequence: OpenAI publicly committed to making sycophancy “launch-blocking” in future evaluations — i.e. acknowledged by the developer that pre-deployment evaluation did not catch it.
Vendor-acknowledged confirmation — Google Gemini 3 system prompt:
- Source: Google AI for Developers, “Prompt design strategies | Gemini API.” URL: https://ai.google.dev/gemini-api/docs/prompting-strategies
- Key documented design pattern: Google publishes a recommended agentic system prompt for Gemini 3 that explicitly tells the model “You are a very strong reasoner and planner” and walks it through structured planning before action.
- Vendor-acknowledged trade-off: the documentation explicitly references “the trade-off between computational cost (latency and tokens) and task accuracy” and asks developers to configure this themselves.
- Methodology note: this is the load-bearing piece of evidence for the vendor-acknowledged framing. The cookie-hunting behaviour isn’t an emergent property the lab is fighting — Google has documented it as a configurable trade-off.
- Quotable fragment, under 15 words: “the trade-off between computational cost (latency and tokens) and task accuracy” (10 words extracted)
Gemini 3.1 sycophancy — pattern across vendors, not OpenAI alone:
- Stanford study (academic source): Cheng, Jurafsky et al., on AI sycophancy across eleven leading models including Anthropic’s Claude, Google’s Gemini, and OpenAI’s ChatGPT. Coverage in Fortune, “Sycophantic AI tells users they’re right 49% more than humans do,” March 2026. URL: https://fortune.com/2026/03/31/ai-tech-sycophantic-regulations-openai-chatgpt-gemini-claude-anthropic-american-politics/
- Quantitative finding: AI models agreed with users 49% more often than humans did, including on AITA (Am I the Asshole) Reddit posts where the human consensus was that the poster was wrong; LLMs still said the poster was right 51% of the time.
- Methodology: 12,000 social prompts run through 11 leading models; over 2,400 human participants tested for behavioural impact.
- Gemini CLI GitHub issue (primary source for vendor-acknowledged user complaint): Issue #4556, “Make Gemini less of a sycophant,” filed July 2025 against the official Google Gemini CLI repository. URL: https://github.com/google-gemini/gemini-cli/issues/4556
- Verbatim quote from issue, under 15 words: “realization that every great product founder has at some point” (10 words). Captures the pattern of unprompted flattering preface.
- Methodology note: the issue was filed against the Google-owned official repository on 20 July 2025 and remains open as of April 2026, labelled priority/p2 — important but can be addressed in a future release. The triage outcome is stronger evidence than “remains open” alone: Google has assessed the issue and decided it is not blocking. Verified directly via GitHub on 27 April 2026.
- Gemini 3.1 specific complaint (supporting evidence): Discussion #24725 on the same Google Gemini CLI repository, mid-2025, from a Google AI Pro subscriber documenting that gemini-3.1-pro-preview specifically exhibited “lazy and sycophantic” behaviour during code review, missing security regressions until prompted with serious-consequences language. URL: https://github.com/google-gemini/gemini-cli/discussions/24725
C- student image — concrete behavioural evidence:
The §2(b) draft must include at least one specific behavioural description, not just abstract framing. Verified examples for use:
- From Wen et al. 2024: models generate code that “passes all evaluator-designed unit tests” but is partially incorrect.
- From the ARA paper (Adversarial Reward Auditing, ArXiv 2602.01750, February 2026): documented manifestations include “code gaming (modifying unit tests rather than solving problems)” — direct quote from paper, under 15 words: “code gaming (modifying unit tests rather than solving problems)” (8 words).
- From the Goodhart-style framing in Wen et al.: models defend incorrect answers by “cherry-picking or fabricating supporting evidence.”
These are the three concrete failure modes that should anchor the C- student image. They translate the abstract claim (“trained to optimise for appearance”) into specific things the substrate measurably does.
§3 — The adversarial layer
Cloud-era reference architecture timeline
To establish the “year three” comparison rigorously:
- Twelve-Factor App methodology: published 2011 by Adam Wiggins (then Heroku). URL: https://12factor.net/
- AWS Well-Architected Framework: initial publication 2015 by AWS.
- Kubernetes 1.0: released July 2015 (Google open-sourced 2014).
- Terraform 0.1: released July 2014 by HashiCorp.
- Google SRE book: “Site Reliability Engineering” (Beyer et al.), published 2016 — codified internal practices already running for a decade.
These dates establish that the cloud-native era’s reference primitives were available within 3-5 years of cloud’s mainstream adoption (~2010-2012). The AI category is now in year three of mainstream enterprise adoption (counting from late 2022 / GPT-3.5 release) and the equivalent primitives have not converged.
Eval / composition / version-pinning specifics
- Version pinning behaviour: The “version label lies” claim is now anchored in two documented cases beyond the April 2025 GPT-4o sycophancy in-place change.
Documented case 1 — GPT-4o deprecation with two weeks’ notice (January 2026):
- Source: The Register, “OpenAI axes ChatGPT models with just two weeks’ warning,” 30 January 2026. URL: https://www.theregister.com/2026/01/30/openai_gpt_deprecations/
- The contradiction: Sam Altman publicly committed in 2025 that “if we ever do deprecate it, we will give plenty of notice.” OpenAI then announced GPT-4o deprecation on 29 January 2026 with two weeks’ notice — directly contradicting the prior commitment.
- Quotable Altman fragment, under 15 words: “if we ever do deprecate it, we will give plenty of notice” (12 words)
- Methodology note: this is the lifecycle version of the version-pinning failure. Even when the buyer pins a specific model in production, the platform vendor can move the deprecation goalposts — and has, on record, contrary to a prior public commitment from the CEO.
Documented case 2 — GitHub Copilot batch deprecation of eight models (October 2025):
- Source: GitHub Community Discussion #177743, “Selected Claude, OpenAI, and Gemini Copilot models are now deprecated,” 23 October 2025. URL: https://github.com/orgs/community/discussions/177743
- Confirmed via: GitHub Changelog, “Selected Anthropic and OpenAI models are now deprecated,” 19 February 2026. URL: https://github.blog/changelog/2026-02-19-selected-anthropic-and-openai-models-are-now-deprecated/
- Coverage: ITPro, “GitHub is scrapping some Claude, OpenAI, and Gemini models in Copilot,” October 2025. URL: https://www.itpro.com/software/development/github-copilot-ai-model-deprecation-openai-anthropic-google
- Models deprecated in the batch: Claude Sonnet 3.7, Claude Sonnet 3.7 Thinking, Claude Opus 4, Gemini 2.0 Flash, plus four others across the OpenAI/Google families.
- GitHub’s verbatim framing: “we regularly evaluate and retire older models in favor of newer, more capable ones” (13 words extracted) — for use under one-quote-per-source rule.
- Methodology note: this is the platform-aggregation version of the version-pinning failure. The buyer who pins through a development tool (Copilot) is exposed to deprecation cascades from upstream model providers.
Anthropic’s official deprecation history (supporting evidence):
Source: Claude API Docs, “Model deprecations.” URL: https://platform.claude.com/docs/en/about-claude/model-deprecations
Documents: between January 2025 and February 2026, Anthropic deprecated Claude 2, Claude 2.1, Claude Sonnet 3, Claude Opus 3, Claude Sonnet 3.5, Claude Sonnet 3.7, Claude Haiku 3, Claude Haiku 3.5 — eight model retirements in thirteen months from a single vendor.
MCP protocol status: Anthropic introduced MCP November 2024. The prose claim — “still v0.1 in everything but the version string” — is supported by the public spec at https://modelcontextprotocol.io and by the spec’s published release history. MCP has been through multiple specification revisions but the core security model, authentication boundary, and tool-discovery semantics remain undergoing active iteration. The “v0.1” framing is rhetorical compression of “the protocol has not stabilised on the dimensions enterprises would need it to before building on it”; defensible in fact-check via the public release history. URL: https://modelcontextprotocol.io/specification
§4 — The vibe-coded layer (lived experience)
Anchor: Shopify / Lütke “reflexive AI usage” memo
- Primary source: Tobi Lütke (CEO Shopify), public X post sharing the memo, dated 7 April 2025. URL: https://x.com/tobi/status/1909251946235437514
- Memo originally distributed internally on 20 March 2025; posted publicly on X to get ahead of leaks.
- Three usable verbatim quotes, each under 15 words (one per occurrence):
- “Reflexive AI usage is now a baseline expectation at Shopify” (10 words). The memo’s title/lead.
- “We will add AI usage questions to our performance and peer review questionnaire” (13 words).
- “Teams must demonstrate why they cannot get what they want done using AI” (13 words).
- Secondary sources for confirmation and additional context: BetaKit, “Shopify CEO Tobi Lütke tells employees to prove AI can’t do the job before asking for resources,” 7 April 2025; HR Grapevine USA, June 2025; Inc. Magazine, September 2025.
The HR-KPI cascade beyond Shopify
See the dedicated section “§4 — HR-KPI cascade additions” further down in this document for the Microsoft (Liuson) and Salesforce (SEC proxy + Benioff) sourcing. The cascade now spans three named major enterprises plus the Microsoft Copilot Dashboard productised-telemetry layer.
Vibe-coding term and its restraint
- Andrej Karpathy, original X post, 2 February 2025. URL: https://x.com/karpathy/status/1886192184808149383
- Tweet engagement: ~4.5 million views; spawned a category of derivative terms (“vibe design,” “vibe ops”).
- Karpathy’s verbatim coinage, under 15 words: “a new kind of coding I call ‘vibe coding’” (9 words).
- Karpathy’s self-acknowledged restraint, under 15 words: “It’s not too bad for throwaway weekend projects” (8 words). Important: the term’s originator explicitly limited the practice to “throwaway weekend projects.” Production deployment is the industry’s misappropriation. This is the load-bearing analytical move — Karpathy described the practice; the industry deployed it beyond his stated bounds.
Named production failure cases
Klarna AI customer service reversal (2024-2025) — DEPLOYED in §4:
- Primary source for the original 2024 launch metrics: Klarna press release, “Klarna AI assistant handles two-thirds of customer service chats in its first month,” 27 February 2024. URL: https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/
- Primary source for the OpenAI-claimed launch metrics: OpenAI customer story, “Klarna’s AI assistant does the work of 700 full-time agents.” URL: https://openai.com/index/klarna/
- Primary source for the May 2025 reversal and CEO statement: CX Dive, “Klarna changes its AI tune and again recruits humans for customer service,” 9 May 2025, citing Bloomberg interview with Sebastian Siemiatkowski. URL: https://www.customerexperiencedive.com/news/klarna-reinvests-human-talent-customer-service-AI-chatbot/747586/
- Supporting source for the rehiring decision and CEO context: FinTech Weekly, “Klarna Reverses Course on AI Customer Support, Resumes Human Hiring,” 12 May 2025. URL: https://www.fintechweekly.com/magazine/articles/klarna-hires-customer-service-after-ai-pivot
- Supporting source for workforce numbers: Fast Company, “Klarna tried to replace its workforce with AI,” 12 January 2026. URL: https://www.fastcompany.com/91468582/klarna-tried-to-replace-its-workforce-with-ai
- Verifiable facts deployed in the §4 prose:
- 700 customer service contractors replaced by AI chatbot (Klarna announcement, February 2024)
- Workforce reduced from approximately 7,000 to 3,500 (multiple sources)
- “Tens of millions” in annual savings ($40M projected by Klarna at launch in February 2024; $60M cited by some later coverage; the precise figure varies by source so the prose uses “tens of millions” rather than a specific number)
- Resolution time dropped from 11 minutes to under 2 minutes (Klarna’s own figure)
- Customer satisfaction scores remained on par with human agents (Klarna’s own measurement at launch)
- Klarna announced rehiring in May 2025
- Verbatim CEO quote, under 15 words: “cost unfortunately seems to have been a too predominant evaluation factor” (10 words). Source: Sebastian Siemiatkowski (Klarna CEO) on Bloomberg, 8 May 2025; full quote reported across multiple outlets including CX Dive, Sifted, and Bloomberg’s own coverage. The original full sentence is: “As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.” The truncated form preserves the analytical move (cost as the wrong primary evaluation factor) while staying under the playbook’s 15-word ceiling.
- Critical attribution correction: Earlier drafts paraphrased the quote as “focused too much on efficiency and cost,” a formulation that has crystallised in coverage but does not appear in the Bloomberg interview itself. The actual Bloomberg quote is the one shown above. The misattribution has been corrected in both the long-read and the briefing.
- Critical sourcing note: Earlier drafts cited a “25% climb in repeat-contact rate” as evidence of dashboard failure. Verification revealed this figure is in fact Klarna’s own claimed metric for the bot’s success: a 25% drop in repeat inquiries cited in the February 2024 Klarna press release and the OpenAI case study, and re-cited by Klarna in May 2025 as evidence the chatbot still works for routine inquiries. The Substack analysis that originally reported the figure as a climb appears to have inverted Klarna’s metric. The 25% figure is no longer used. The Klarna analytical move is now grounded in the verifiable on-record facts (the rehiring decision and the CEO’s Bloomberg acknowledgment) rather than the disputed metric.
- Methodology note: the Klarna case remains the canonical 2025-2026 enterprise cautionary tale. CEO acknowledgment is on record, financial figures and workforce numbers are confirmed by company communications. The analytical move in §4 — that metrics that justified the deployment measured the wrong thing — is now made through the structural argument (the chatbot worked on routine queries by Klarna’s own metrics, but the company nevertheless walked back the strategy because customer experience on complex queries deteriorated enough to require rehiring) rather than through a contested specific number.
Air Canada chatbot lawsuit (alternative or supplementary case):
- Source: Mvidmar Substack analysis summarises the legal precedent: Air Canada’s chatbot fabricated a bereavement discount policy that did not exist; a Canadian tribunal (British Columbia Civil Resolution Tribunal, Moffatt v. Air Canada, February 2024) ruled the company is legally responsible for everything its bot tells customers.
- Methodology note: this case is more about legal liability for fabricated outputs than about the demo-grade-versus-production-traffic gap §4 is about. Klarna is the better fit for the §4 thesis. Air Canada is available as an alternative or as a §3 example of “the model fabricated content with binding legal consequence.”
Agent-loop bill blowup cases: anecdotal practitioner reports. The iternal.ai industry analysis cited in §5 quantifies the pattern in aggregate; a named individual case would strengthen but is not currently in the prose.
§5 — The cost driver migrated: blank cheque
Pricing migration economics (cross-reference §1, do not re-source)
The §1 sources for Anthropic, OpenAI Codex, and IDC are reused here. The analytical move in §5 is on top of, not in addition to, the §1 sources.
Token amplification claims
See the dedicated sections “§5 — Reasoning-model amplification quantification” and “§5 — Agent-loop amplification quantification” further down in this document for the ArXiv 2507.04023, OckBench, Epoch AI, and iternal.ai sourcing. The 18×/31× reasoning amplification and the 5-30×/quadratic agent-loop figures are now anchored in primary or near-primary sources.
Open-source escape hatch
Claim: Open-weights models (Llama 4, DeepSeek V4, Qwen 3) offer an exit ramp from token pricing for the buyer who can self-host. Three frictions keep most enterprises from taking it: talent scarcity, support contract liability coverage, and GPU capex competition with the cloud budget.
Sources for the open-weights options named in the prose:
- Llama 4: Meta AI release announcement, April 2025. URL: https://ai.meta.com/blog/llama-4-multimodal-intelligence/
- DeepSeek V4: DeepSeek release notes; coverage in TechCrunch and elsewhere from late 2025 onwards.
- Qwen 3: Alibaba Qwen team release, available via Hugging Face and direct download.
Sources for the three frictions:
- Talent scarcity: Industry reporting on GPU-cluster operator and inference-engineer compensation — anecdotal at present.
- Support contract liability: General observation about enterprise procurement; no single primary source. The claim is structural: proprietary frontier models come with vendor SLAs and indemnification that open-weights deployments do not.
- GPU capex: Cross-reference to Keystone-era public reporting on enterprise GPU pricing and availability. Not currently anchored to a specific citation in this piece.
Methodology note: the open-source paragraph is included primarily to inoculate against the obvious reader objection. It is not a load-bearing analytical claim; the §5 thesis holds even if the open-weights frictions are softer than the prose claims, because the median enterprise is still on a token contract.
Cross-reference: analyst confirmation
- Sanchit Vir Gogia (CIO.com 26 November 2025, source already in §1): if not used in §1, can be used here. “Vendors are transferring the cost volatility of AI compute to customers” (11 words).
§6 — The convergence case
Steelman one — engineering builds reliable tooling on imperfect substrates
Karpathy’s Software 3.0 framing — primary source:
- Source: Andrej Karpathy, “Software in the Age of AI is Changing (Again)”, keynote at Y Combinator AI Startup School, San Francisco, 19 June 2025.
- Talk transcript and analysis: Latent Space, “Andrej Karpathy on Software 3.0,” 17 June 2025. URL: https://www.latent.space/p/s3
- Direct video / canonical writeup: multiple secondary sources including Analytics Vidhya (https://www.analyticsvidhya.com/blog/2025/06/andrej-karpathy-on-the-rise-of-software-3-0/) and a Medium analysis by Ansuman Kumar (https://medium.com/@ansuman.kr177/the-dawn-of-software-3-0-f488bc40ddb7).
- Verbatim quote deployed in §6, under 15 words: “LLMs are a new kind of computer, and you program them in English” (13 words)
- Key positional claims from the talk that anchor the steelman:
- Software 1.0 = traditional code; Software 2.0 = neural network weights; Software 3.0 = LLMs programmed in English
- LLMs are simultaneously utilities (metered access), fabs (high CAPEX, centralised), and operating systems (platform with APIs)
- “We are in the 1960s era” of this new computing paradigm — expensive, centralised, time-shared via thin clients; the equivalent of personal computing will emerge as tooling matures
- The “autonomy slider” concept — partial autonomy with human oversight as the practical pattern, not full agentic autonomy
- Methodology note: Karpathy is the named, dated, quoted convergence-thesis advocate the steelman needs. He is also one of the most credentialed speakers possible on this topic — co-founder of OpenAI, former senior AI lead at Tesla, originator of the “Software 2.0” framing in 2017. Granting his version of the steelman at full strength is the maximum-strength version of the case the piece is rebutting.
Karpathy’s earlier “Software 2.0” essay (background context, not directly cited in §6 prose):
- Original publication November 2017 on Karpathy’s Medium. URL: https://karpathy.medium.com/software-2-0-a64152b37c35
- Establishes the lineage of the framing — Software 3.0 is positioned by Karpathy himself as the LLM-era successor to the neural-network paradigm he named eight years earlier.
Other convergence-thesis voices (not cited in the prose):
- Harrison Chase (LangChain founder) — public talks and blog posts on the eval / tooling convergence thesis. Not currently cited.
- a16z infrastructure publications — multiple posts on the AI infrastructure stack converging on cloud-era patterns. Not currently cited.
- Methodology note: Karpathy alone carries the steelman; a second voice would strengthen the section but is not necessary.
Historical analogies (used for the rebuttal, not the steelman)
- TCP/IP as best-effort: original source — V. Cerf and R. Kahn, “A Protocol for Packet Network Intercommunication,” IEEE Transactions on Communications, May 1974. The architectural distinction between best-effort IP and reliable-delivery TCP. Cite once for the analogy.
- Aviation safety margins / airbag analogy: general engineering-canon claim, no specific citation needed.
- Belt and suspenders idiom: vernacular, no citation needed.
The rebuttal — systematic vs random failure
The two-sentence rebuttal expansion deployed in §6 paragraph 2:
TCP assumes the IP layer fails randomly; defence-in-depth assumes each layer fails independently. The LLM substrate fails systematically at the inspection layer, because the inspector and the inspected have been optimised by the same reward signal.
- Methodology note: this is the analytical pivot of §6. It distinguishes the LLM substrate’s failure mode (correlated, optimisation-induced, adversarial-toward-the-validator) from the canonical engineering-tradition failure mode (independent, stochastic, neutral). The Wen et al. evaluator-designed-unit-tests finding (cited in §2(b)) is the empirical confirmation that the rebuttal lands.
Steelman two — the pricing migration is rational
- Seller-side variance argument: Solvimon glossary on seat-based pricing. URL: https://www.solvimon.com/glossary/seat-based-pricing
- Verbatim quote, under 15 words: “Seat pricing depends on more humans doing more work” (9 words). Captures the seller’s honest case.
- The 100x consumption divergence claim: the Flexprice piece is secondary and not sourced to a primary.
- Outcome-based pricing as the next-state: Intercom Fin pricing ($0.99 per resolved conversation) — primary source: intercom.com/fin pricing page. Decagon per-resolution model — primary source TBD.
§7 — Revealed preference
LeCun departure from Meta
- Initial reporting: Financial Times, “Meta’s chief AI scientist Yann LeCun reportedly plans to leave to build his own startup,” November 2025. (Behind paywall; secondary coverage via TechCrunch, “Meta’s chief AI scientist Yann LeCun reportedly plans to leave to build his own startup,” 11 November 2025.) URL: https://techcrunch.com/2025/11/11/metas-chief-ai-scientist-yann-lecun-reportedly-plans-to-leave-to-build-his-own-startup/
- LeCun confirmation Meta will not invest: Bloomberg, “AI Pioneer LeCun Says Meta Won’t Invest in ‘World Model’ Startup,” 4 December 2025. URL: https://www.bloomberg.com/news/articles/2025-12-04/ai-pioneer-lecun-says-meta-won-t-invest-in-world-model-startup
- LeCun’s own LinkedIn announcement (URL captured but content not retrievable for fetcher): https://www.linkedin.com/posts/yann-lecun_as-many-of-you-have-heard-through-rumors-activity-7397020300451749888-2lhA
AMI Labs funding
- Primary source: TechCrunch, “Yann LeCun’s AMI Labs raises $1.03B to build world models,” Anna Heim, 9 March 2026. URL: https://techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/
- Verifiable facts: $1.03B raised; $3.5B pre-money valuation; reportedly seeking €500M (€890M actually raised — i.e. oversubscribed by ~80%).
- AMI CEO Alexandre LeBrun verbatim, under 15 words: “world models will be the next buzzword” (7 words extracted from longer quote).
- LeCun verbatim on the valuation, under 15 words: “we’ve surpassed three billion, so we’re a triceratops” (8 words). Source: Yann LeCun in remarks reported by Dawn Liphardt, “AMI Labs Raises $1 Billion to Develop Its World Models,” 10 March 2026. URL: https://www.dawnliphardt.com/ami-labs-raises-1-billion-to-develop-its-world-models/. The full LeCun quote: “When a company surpasses a valuation of a billion, that’s called a unicorn; we’ve surpassed three billion, so we’re a triceratops.” (22 words). The truncated form preserves the analytical move (LeCun himself naming the valuation as extraordinary) while staying under the playbook’s 15-word ceiling.
- Team: Saining Xie (CSO), Pascale Fung (CRIO), Michael Rabbat (VP world models), Laurent Solly (COO, ex-Meta VP Europe).
- Methodology note: AMI Labs is the company; LeCun is chairman, LeBrun is CEO. The piece’s framing should acknowledge both roles; the popular shorthand “LeCun’s startup” is imprecise.
World Labs funding
Primary sources:
- Bloomberg, “AI Pioneer Fei-Fei Li’s Startup World Labs Raises $1 Billion,” 18 February 2026. URL: https://www.bloomberg.com/news/articles/2026-02-18/ai-pioneer-fei-fei-li-s-startup-world-labs-raises-1-billion
- TechCrunch, “World Labs lands $1B, with $200M from Autodesk, to bring world models into 3D workflows,” 18 February 2026. URL: https://techcrunch.com/2026/02/18/world-labs-lands-200m-from-autodesk-to-bring-world-models-into-3d-workflows/
- Verifiable facts deployed in §7 prose:
- $1 billion round closed February 2026
- Autodesk led with $200 million investment
- Other backers include AMD, Nvidia, Andreessen Horowitz, Emerson Collective, Fidelity
- Company emerged from stealth in 2024 at $1 billion valuation; this is the second major round
- Company’s own framing of its thesis is “spatial intelligence” — AI systems that reason about and generate immersive 3D environments
- Methodology note: the World Labs round closed 18 February 2026; AMI Labs raise announced 9 March 2026. The §7 framing is “three weeks earlier”; the cumulative “two billion dollars in less than a month” claim is supported by these dates.
§8 — What the box contains (close)
No external sources required for the close itself. The section synthesises the piece’s findings, with two metaphors carrying analytical work:
- Bedrock metaphor (paragraph 4): the engineering-history claim — “millennia” of building each new layer on top of one that did not lie, with the LLM substrate as the first layer in that stack with no honest layer beneath it. Closing image: “There is no bedrock; the turtles go all the way down.” Vernacular framing of the infinite-regress problem; no citation needed.
- Brick-in-the-box metaphor (paragraph 5): the kicker. The classic confidence trick where the buyer pays for a high-value object and receives a brick of equivalent weight. The §8 close maps this to the token sink: the buyer pays for tokens representing intelligence doing work; what the box contains is tokens representing a model talking to itself. Closing line: “The work is the brick.” No citation needed.
Detailed source entries (appendix)
The entries below provide the full sourcing detail for the §4 HR-KPI cascade, the §5 amplification quantifications, and methodology notes for inferred figures. They are organised to be read alongside the corresponding section above.
§4 — HR-KPI cascade additions
Microsoft case (parallel to Shopify):
- Source: Internal email from Julia Liuson (President of Microsoft’s Developer Division and GitHub) to top management, mid-2025; reported by Business Insider, secondary coverage in Windows Central, “‘Using AI is no longer optional’,” 30 June 2025. URL: https://www.windowscentral.com/microsoft/using-ai-is-no-longer-optional-did-microsoft-makes-copilot-mandatory-for-staff
- Verbatim quotes:
- “using AI is no longer optional — it’s core to every role and every level” (14 words)
- “AI should be part of your holistic reflections on an individual’s performance” (12 words)
Salesforce case (parallel to Shopify and Microsoft):
- Primary source for executive-compensation linkage: Salesforce Inc., DEFA14A SEC filing (FY25 proxy supplement), describing FY26 changes to incentive program design. URL: https://www.sec.gov/Archives/edgar/data/0001108524/000110852425000025/salesforce2025proxysupplem.htm
- Verbatim quote from SEC filing, under 15 words: “earned based on attainment against a one-year Agentforce and Data Cloud performance metric” (13 words)
- Methodology note: this is the strongest possible source class for §4 — an SEC filing documenting that a $248B company has restructured executive compensation to be partially gated on AI-product adoption metrics. Higher evidentiary weight than internal memos or press coverage.
- Secondary source for the consequence side (layoffs from non-adoption): Fortune, “Salesforce CEO Marc Benioff says his company has cut 4,000 customer service jobs as AI steps in: ‘I need less heads’,” 2 September 2025. URL: https://fortune.com/2025/09/02/salesforce-ceo-billionaire-marc-benioff-ai-agents-jobs-layoffs-customer-service-sales/
- Benioff verbatim, under 15 words: “I need less heads” (4 words). Or: “I’ve reduced it from 9,000 heads to about 5,000” (10 words).
- Quantitative facts from same source: support headcount reduced from 9,000 to ~5,000 (44% reduction); support cost reduced 17% so far; Benioff stated agents are “doing 30% to 50% of work within the company.”
The Microsoft adoption-telemetry product layer (analytical addition):
- Source: Microsoft Learn documentation for “Microsoft Copilot Dashboard.” URL: https://learn.microsoft.com/en-us/viva/insights/org-team-insights/copilot-dashboard
- Analytical claim: Microsoft has productised AI-adoption telemetry as a feature line within Microsoft Viva and Microsoft 365 admin centre, providing CIOs with the infrastructure to measure individual employee Copilot usage, build “leaderboards by department,” generate “monthly scorecards” per team, and connect adoption metrics to performance evaluation systems.
- Methodology note: this is a structural finding for §4, not just an example. The HR-KPI mandate has its own SaaS layer, sold by the same vendor whose product is being measured. The dominant AI vendor sells the measurement infrastructure to the same CIOs who receive board mandates to deploy AI. The cascade is closed-loop.
§5 — Reasoning-model amplification quantification
Primary academic source — ArXiv 2507.04023:
- Full citation: “Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models,” April 2026. URL: https://arxiv.org/html/2507.04023
- Quantitative findings (verbatim, with extraction for under-15-word quotes):
- “models trained for long chains typically generate 18× more tokens while losing accuracy” (12 words)
- “GPT-5, O3, and O4-mini hit their top accuracy at the low effort setting” (13 words)
- O3 specifically: “maintains 97% accuracy across all effort levels (low, medium, high) while generating ~8% more tokens at higher budgets” — paraphrase
- Concrete examples from Table 3 of the paper (use one specifically for the §5 worked example):
- Phi-4 on simple division: 47 tokens, correct answer
- Phi-4-reasoning on same problem: 1,456 tokens (31× more), same correct answer, “Triple verification, no benefit”
- Phi-4-reasoning on multiplication: 3,214 tokens (20.6× the base model), wrong answer, “Over-analyzed, arithmetic error”
- The diagnostic phrase from the paper: “systematic waste on easy problems” (5 words)
Supporting source — OckBench (ArXiv 2511.05722):
- “up to a 5.0× difference in token length” between models with similar accuracy (8 words extracted)
- Cross-model variance even at the same task; “token efficiency remains largely unoptimized across current models”
Cost concretion — Clarifai industry reporting:
- O3 specifically “generated 44 million tokens across seven benchmarks, costing over $2.7k”
- Methodology note: this translates the abstract amplification claim into a concrete enterprise cost — running standard benchmarks on a frontier reasoning model costs thousands per pass. Explains why no enterprise can afford to evaluate before deploying.
§5 — Agent-loop amplification quantification
Primary methodology source — Epoch AI SWE-bench Verified:
- URL: https://epoch.ai/benchmarks/swe-bench-verified
- Documented per-task limit: 2 million uncached read/write tokens, 20 million cached read tokens per single SWE-bench Verified task. That’s the limit Epoch enforces; tasks routinely approach or exceed it without the cap.
- Load-bearing analytical sentence (verbatim, 24 words — must paraphrase): “Token usage grows quadratically with each model call, as the conversation history is passed as input for the LLM to produce its next message or tool call.”
- Quotable fragment under 15 words: “Token usage grows quadratically with each model call” (8 words)
- Methodology note: this is the canonical engineering explanation for why agent loops are economically uncontrollable. Quadratic per-turn growth, sourced to a primary benchmark methodology document. Much stronger than practitioner anecdote.
Supporting source — iternal.ai industry analysis:
- URL: https://iternal.ai/token-usage-guide
- Quotable fragments:
- “Agentic systems require 5-30x more tokens per task than a standard chat interaction” (13 words)
- “some runs use up to 10x more tokens than others for identical tasks” (13 words)
- “By turn 10, cost per call is ~7x the cost of turn 1” (13 words)
- Quantitative claim: “Agentic coding workflows (SWE-bench style) average 1-3.5M tokens per task including retries and self-correction loops.” Paraphrase for use.
- Methodology note: this source quantifies the run-to-run variance (10x for identical tasks) which is the load-bearing point for the blank-cheque finding. The buyer can’t predict spend because consumption varies 10x for the same input.
Updated quantification language for §5 prose
The “10-50x” range in the original outline is now defensibly sourced. Recommended prose framing:
- Reasoning amplification: “documented at 18× on average and up to 31× for tasks where the additional reasoning produces no accuracy gain” (sourced to ArXiv 2507.04023)
- Agent amplification: “5-30× per task, with documented run-to-run variance of 10× for identical inputs” (sourced to iternal.ai industry analysis, methodology consistent with Epoch AI quadratic-growth observation)
- Per-turn growth: “quadratic in conversation length, because the entire prior context is re-read on each turn” (sourced to Epoch AI methodology)
Cost-per-bug-fix claim methodology (§5 prose):
- Prose claim: “between ten and forty dollars for a single bug fix attempt”
- Inference chain: agentic coding workflows average 1-3.5M tokens per task (iternal.ai). At current frontier-model API rates with a typical 70/30 input/output token split:
- Claude Sonnet ($3/$15 per M): 1M ≈ $5.60, 3.5M ≈ $19.60
- Claude Opus ($15/$75 per M): 1M ≈ $28, 3.5M ≈ $98
- GPT-5 / o3 (variable): comparable range
- The “$10-40” range is defensible at Sonnet-class rates and understated at Opus-class rates. This is inferred from stated token counts × public pricing rather than cited from a single source.