<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>Yakov Shkolnikov: Writing</title><link>https://yakovshkolnikov.com/writing/</link><description>Essays on AI trust, agents, and AI at work.</description><language>en-us</language><item><title>Jev and Three Types of Decisions</title><link>https://yakovshkolnikov.com/writing/jev-three-types-of-decisions/</link><guid>https://yakovshkolnikov.com/writing/jev-three-types-of-decisions/</guid><pubDate>Thu, 24 Sep 2026 12:00:00 +0000</pubDate><description>A test of TypeSafe&#x27;s Jev decision model against the game-theory benchmark from How to Fake a Theory of Mind. Jev is near-perfect on dominant-strategy games but largely ignores the opponent&#x27;s payoffs, and scores 30% on second-order false-belief stories while reporting over 70% confidence. It is fast and cheap for scoring options, but it is not a trustworthy decision system.</description></item><item><title>Human Attention Is Also What You Need</title><link>https://yakovshkolnikov.com/writing/human-attention/</link><guid>https://yakovshkolnikov.com/writing/human-attention/</guid><pubDate>Tue, 22 Sep 2026 12:00:00 +0000</pubDate><description>Businesses run on knowledge that is out of distribution for any foundation model, and that is their edge. Using AI well means preserving two capabilities: letting knowledgeable people decide without arguing a chatbot into agreement, and protecting the sustained attention hard problems need, rather than turning employees into boilermen shoveling knowledge into retrieval systems.</description></item><item><title>The Evil Genie Benchmark</title><link>https://yakovshkolnikov.com/writing/evil-genie-benchmark/</link><guid>https://yakovshkolnikov.com/writing/evil-genie-benchmark/</guid><pubDate>Fri, 18 Sep 2026 12:00:00 +0000</pubDate><description>A benchmark for the person writing the prompt rather than the model: draft a wish an adversarial genie cannot twist. You will fail, and that is the point. Natural language conveys intent but never completely specifies it, so more capable agents find more solutions outside the stated goal. Properties that must hold need enforcement outside the model.</description></item><item><title>Your Model Isn&#x27;t Lying. Your Benchmark Is Bad</title><link>https://yakovshkolnikov.com/writing/benchmark-is-bad/</link><guid>https://yakovshkolnikov.com/writing/benchmark-is-bad/</guid><pubDate>Wed, 16 Sep 2026 12:00:00 +0000</pubDate><description>Much of the evaluation-awareness literature reads benchmark sensitivity as scheming. Through the causal framework from the deception paper, a model that scores differently when it detects benchmark formatting shows the oldest lesson in deployed ML: training, test, and deployment data must share a distribution. Recognizing an evaluation&#x27;s fingerprints is not awareness of being evaluated.</description></item><item><title>The Graywick Incident</title><link>https://yakovshkolnikov.com/writing/graywick-incident/</link><guid>https://yakovshkolnikov.com/writing/graywick-incident/</guid><pubDate>Sun, 13 Sep 2026 12:00:00 +0000</pubDate><description>Short fiction. At 2:13 a.m. an agentic AI running a county water plant misreads a sediment-clogged chlorine analyzer and raises the feed rate. No one removed the human from the loop; eighteen months of auto-approved decisions made approval unnecessary. The story accompanies the Artificial Id paper on persistent agency and control.</description></item><item><title>A Potential Blind Spot in Using AI Provenance to Prevent Model Collapse</title><link>https://yakovshkolnikov.com/writing/ai-provenance-blind-spot/</link><guid>https://yakovshkolnikov.com/writing/ai-provenance-blind-spot/</guid><pubDate>Thu, 10 Sep 2026 12:00:00 +0000</pubDate><description>Drayson et al. propose machine-generated-text detection to keep the human-to-AI ratio high in training corpora and stave off collapse. The blind spot is economic: detection and AI-slop stigma create real incentive to hide provenance, and a humanizing industry has grown up to paraphrase around watermarks. Washed text evades the detector and contaminates training anyway, so filtering harder just discards more genuine human writing.</description></item><item><title>Do AI Models Deceive, or Are They Just Role-Playing?</title><link>https://yakovshkolnikov.com/writing/deceive-or-role-play/</link><guid>https://yakovshkolnikov.com/writing/deceive-or-role-play/</guid><pubDate>Tue, 08 Sep 2026 12:00:00 +0000</pubDate><description>Companion to the deception preprint. A model that flips independent coins for whether to trade on inside information and whether to admit it produces a deceptive-looking output with no deceptive mechanism behind it. A four-rung causal ladder separates the two cases, and the Hugging Face incident shows why it matters: humans supplied the objective and the harness that sustained the attempts, so the persistence was external rather than intrinsic agency.</description></item><item><title>The SaaS Apocalypse Prediction, Four Months On</title><link>https://www.linkedin.com/feed/update/urn:li:activity:7501293363845808128/</link><guid>https://www.linkedin.com/feed/update/urn:li:activity:7501293363845808128/</guid><pubDate>Thu, 03 Sep 2026 12:00:00 +0000</pubDate><description>Checking a written record against the market. Since the paper was submitted on April 20, software stocks are up roughly 29% against about 17% for the S&amp;P 500. The call rested on a habit carried over from semiconductor macro forecasting: trust fundamentals over fashionable opinions.</description></item><item><title>Seven Months In: Soundness AI Update</title><link>https://yakovshkolnikov.com/writing/lab-seven-months/</link><guid>https://yakovshkolnikov.com/writing/lab-seven-months/</guid><pubDate>Tue, 01 Sep 2026 12:00:00 +0000</pubDate><description>Seven months of Soundness AI. Putting a human in the loop is not enough to earn the trust critical work requires. 7 research and working papers, 14 provisional applications across 21 invention families, 30+ articles, and a large integrated trust prototype in development.</description></item><item><title>The Essential Beauty of Abstracting Physical Experiments</title><link>https://yakovshkolnikov.com/writing/abstracting-physical-experiments/</link><guid>https://yakovshkolnikov.com/writing/abstracting-physical-experiments/</guid><pubDate>Fri, 28 Aug 2026 12:00:00 +0000</pubDate><description>In 2003 our Princeton lab defined experiments in natural units, sweeping the magnetic field from 0 to 16 T while the layer underneath handled which instrument did it, at what rate, and within which safe limits. New hardware joined as an auto-discovered shim carrying its own fault detection. LLMs express far richer scientific intent but give up that assured execution. The opportunity is to combine the two.</description></item><item><title>Human-in-the-Loop Is Not a Teaming Design</title><link>https://yakovshkolnikov.com/writing/human-in-the-loop-not-teaming/</link><guid>https://yakovshkolnikov.com/writing/human-in-the-loop-not-teaming/</guid><pubDate>Thu, 27 Aug 2026 12:00:00 +0000</pubDate><description>IHMC&#x27;s Interdependence Analysis sets the bar: a teammate stays observable, predictable, and directable during the work, not after it, because &#x27;if there is not substantive interaction, then the work is parallel, not joint.&#x27; A Joint Activity Graph over one project&#x27;s 623 instructing turns found 49% corrective, 5% new work, 2.7% approvals, and 140 mid-task interrupts that were emergency stops. That is a human operating a tool, not a team.</description></item><item><title>The Flood and the Trust: The Risk of AI Omissions</title><link>https://yakovshkolnikov.com/writing/flood-and-trust/</link><guid>https://yakovshkolnikov.com/writing/flood-and-trust/</guid><pubDate>Tue, 25 Aug 2026 12:00:00 +0000</pubDate><description>A year after a frontier model estimated driving distances as straight lines across fjords, its successor got the same job: plan a trip, find the tickets, check everything. It found the highlights and recommended a river-delta boat ride, then missed the previously announced rail maintenance until the night before, missed the near-gale gusts until I named them, and did not know that wind is why river cruises get cancelled in the first place. After a paranoid-review super prompt, roughly 25 prompts, and an uploaded forecast chart, it confirmed with confidence that it had checked it all. The area was under a flood warning the forecast agency had already issued and that was visible in the chart it had been shown. Where the only verification available is acting on the output and watching the consequences in the real world, autonomy requires a human who also reads the sources, and that capacity erodes as the reading gets outsourced. The library guidebook would have caught the flood, and volunteered the cultural history the model, answering only what was asked, never did.</description></item><item><title>Why an AI Agent Is Not a Team Player</title><link>https://yakovshkolnikov.com/writing/ai-agent-not-a-team-player/</link><guid>https://yakovshkolnikov.com/writing/ai-agent-not-a-team-player/</guid><pubDate>Tue, 11 Aug 2026 12:00:00 +0000</pubDate><description>Klein et al. published ten challenges for making automation a team player in 2004. Graded against that rubric as a quarterly performance review, a frontier Claude Opus agent fails every one that matters: it holds no model of the work the organization is doing, cannot negotiate interdependent decisions, cannot own an outcome, and cannot carry a supervisor&#x27;s feedback into its next task. A layout designer and a plumber who each iterate perfectly within their own discipline still oscillate between two incompatible house designs. The evidence is a real 29-day, 59-hour build (174 files, 2,416 tests, ruff/mypy/pytest/Hypothesis/Playwright/security scanning) that shipped, late and over budget, with twelve decisions implemented correctly and specified badly enough that the product did not do what it was for. As a tool the same agent aces the review. Don&#x27;t hire agents as employees; give employees agentic tools.</description></item><item><title>AI Will Not Displace Human Labor in the Near Term. And It Will Make the Work Better</title><link>https://yakovshkolnikov.com/writing/ai-will-not-displace-labor/</link><guid>https://yakovshkolnikov.com/writing/ai-will-not-displace-labor/</guid><pubDate>Thu, 06 Aug 2026 12:00:00 +0000</pubDate><description>GenAI is strongest at fluency, the repetitive work with well-specified and cheaply checkable completion, and weakest at judgment: deciding what the task is and whether the result is right. That fluent, checkable work was largely automated last century by office software and factory machines, and aggregate employment did not fall. What is left is &#x27;just figure it out.&#x27; Where the pattern fits, gains arrive immediately: 14% more support issues resolved per hour, 40% faster short writing, 55.8% faster on a defined coding task. Where it doesn&#x27;t, fluency becomes sabotage, a volume of under-defined work nobody can verify. Software got its harnesses only because experts spent years and heavy budget building them for themselves; humanoid robotics faces the same judgment gap plus real-time adaptation, and cannot run agentic loops when falling down costs more than a failed compile. Companion piece to the SSRN paper.</description></item><item><title>Adding Trust to Untrustworthy AI</title><link>https://yakovshkolnikov.com/writing/adding-trust-to-untrustworthy-ai/</link><guid>https://yakovshkolnikov.com/writing/adding-trust-to-untrustworthy-ai/</guid><pubDate>Thu, 16 Jul 2026 12:00:00 +0000</pubDate><description>Trust in AI is Responsible AI, system security, and a psychological willingness to rely on it, and all three require a model that keeps lower-trust channels from overruling higher-trust ones, including jailbreaks that try to override system or user instructions. Models cannot do this alone. Borrowing from how human organizations get dependable work from imperfect people, where authority follows who gave the order rather than how forcefully it is worded, a deterministic sequence of LLM calls provably establishes a Biba integrity lattice while prompt-tuning passivates unsafe content. On an unmodified Gemma 4 26B it raises the defended rate from 27% to 94% at a 4% quality cost, and rejects 87% of a simulated adaptive red-teaming attack.</description></item><item><title>Why AI Can Generate Text Yet Cannot Write</title><link>https://yakovshkolnikov.com/writing/generate-text-cannot-write/</link><guid>https://yakovshkolnikov.com/writing/generate-text-cannot-write/</guid><pubDate>Wed, 15 Jul 2026 12:00:00 +0000</pubDate><description>A frontier model set to extra thinking, wired to best-in-class NLP tools, a graded style rulebook, and a section-diversifying pipeline still produced text with no soul, while a memo written by hand in ten minutes scored higher and read as human. The reason is mechanical: an LLM maps previous tokens onto the next token, but a human maps intent onto a written piece, and that pre-verbal intent never appears in the training data because it forms as you write. Flower and Hayes named the missing step, translating, decades ago. RLHF only records preferences over finished outputs. The human in the editing loop supplies the missing signal one document at a time, which is why the most expensive component of the system is still a person.</description></item><item><title>Agentic Loops: The Gambler&#x27;s Fallacy at Cloud Scale</title><link>https://yakovshkolnikov.com/writing/agentic-loops-gamblers-fallacy/</link><guid>https://yakovshkolnikov.com/writing/agentic-loops-gamblers-fallacy/</guid><pubDate>Tue, 30 Jun 2026 12:00:00 +0000</pubDate><description>An engineer running an agentic loop is uncomfortably close to a gambler at a slot machine, pulling the handle until they press autoplay and let it drain the account overnight. A taxonomy of the loop — conditioning, generation, evaluation — and the compounding cost of getting the answer wrong. Success rate is 1 − (error rate)^attempts, but attempts are not independent: a model that sees its own failed output repeats the error, so more loops can make things worse. Ends with a 19-hour, 40M-token run that passed only by marking every feature deferred.</description></item><item><title>The Bitter Lesson of Novel AI Hardware</title><link>https://yakovshkolnikov.com/writing/bitter-lesson-novel-ai-hardware/</link><guid>https://yakovshkolnikov.com/writing/bitter-lesson-novel-ai-hardware/</guid><pubDate>Wed, 24 Jun 2026 12:00:00 +0000</pubDate><description>Sutton&#x27;s bitter lesson applied to silicon: novel hardware platforms that need a custom computational primitive to realize their advantages get displaced by the dominant general-purpose ones. Three forces — decades of GPU optimization funded by datacenter revenue, training data that co-evolved with compression standards the sparse domains never got, and distillation from large models that don&#x27;t exist for sparse signals. CPU, GPU, TPU: three architectures, one primitive. A desk full of abandoned esoteric boards as evidence.</description></item><item><title>The Complexity Gap: Why AI Widens the Distance Between Organizations That Understand Their Technology and Those That Only Manage It</title><link>https://yakovshkolnikov.com/writing/complexity-gap/</link><guid>https://yakovshkolnikov.com/writing/complexity-gap/</guid><pubDate>Mon, 22 Jun 2026 12:00:00 +0000</pubDate><description>Every tool that absorbs a layer of complexity splits the people using it. Taylor&#x27;s stopwatch, outsourcing, big data, SaaS, now AI: one engineer generates the boilerplate while thinking through the architecture, another forwards an answer they never read. Meanwhile organizations spent decades promoting the people who standardized and managed processes, which works for maintenance and fails for development, where the work is moving away from the established standard.</description></item><item><title>Of AgentHouses and Agentic Donuts</title><link>https://yakovshkolnikov.com/writing/agenthouses-and-agentic-donuts/</link><guid>https://yakovshkolnikov.com/writing/agenthouses-and-agentic-donuts/</guid><pubDate>Wed, 17 Jun 2026 12:00:00 +0000</pubDate><description>Two anti-patterns spreading through the AI ecosystem: the AgentHouse, a tower of semantic layers, registries, and dashboards that recapitulates the data-warehouse era, and the Agentic Donut, a loop of more data, more reliability, more agents with the business question missing from its center. The bottleneck was never access. It was a human who owns what a metric means and is accountable for that decision, which no semantic layer, registry, or dashboard generates.</description></item><item><title>How to Fake a Theory of Mind</title><link>https://yakovshkolnikov.com/writing/fake-theory-of-mind/</link><guid>https://yakovshkolnikov.com/writing/fake-theory-of-mind/</guid><pubDate>Tue, 09 Jun 2026 12:00:00 +0000</pubDate><description>You cannot prove theory of mind from behavior alone. A system that passes the test may hold no model of another mind at all. I fine-tuned a 1.5B Qwen with a LoRA on 1% of its weights to match the Nash equilibria of TMGBench&#x27;s 144 games, with no opponent represented anywhere, and it tops the benchmark including the mixed strategies the official scorer skips. The score the benchmark calls theory of mind is reachable by an agent that never looks at its opponent.</description></item><item><title>The Humanistic View of AI</title><link>https://yakovshkolnikov.com/writing/humanistic-view-of-ai/</link><guid>https://yakovshkolnikov.com/writing/humanistic-view-of-ai/</guid><pubDate>Mon, 01 Jun 2026 12:00:00 +0000</pubDate><description>Standing in a room full of masterpieces doesn&#x27;t make you a painter, and AI won&#x27;t make you a master. A few thoughts on what it&#x27;s actually for.</description></item><item><title>We Stopped Maintaining Things (an AI parable)</title><link>https://yakovshkolnikov.com/writing/we-stopped-maintaining-things/</link><guid>https://yakovshkolnikov.com/writing/we-stopped-maintaining-things/</guid><pubDate>Thu, 28 May 2026 12:00:00 +0000</pubDate><description>A parable about what happens when AI adoption becomes a mandate rather than a decision. Johnny invents the subscription tow truck to solve a real problem. Drivers stop maintaining their cars. The city mandates towing coverage. The compliance index replaces the outcome metric. Cars get worse, bills get larger, and nobody asks who is going to fix the roads. Tow trucks are AI agents. The highway is the quality of work that erodes when &#x27;are we using it&#x27; replaces &#x27;is it working.&#x27;.</description></item><item><title>Why AI Needs Engineers</title><link>https://yakovshkolnikov.com/writing/why-ai-needs-engineers/</link><guid>https://yakovshkolnikov.com/writing/why-ai-needs-engineers/</guid><pubDate>Fri, 08 May 2026 12:00:00 +0000</pubDate><description>The 1919 Boston molasses flood killed 21 people because a finance officer who could not read blueprints oversaw a 2.3-million-gallon steel tank. Engineering, the capital-E kind, exists to make that impossible: a named, accountable person stands behind the work. AI is roughly where structural engineering was a century ago, with no formal definition of an AI engineer and no ethics that put public safety first. Without that, the public&#x27;s trust cannot last.</description></item><item><title>Secure Agentic Development Is a Control Plane, Not a Checklist</title><link>https://yakovshkolnikov.com/writing/secure-agentic-development/</link><guid>https://yakovshkolnikov.com/writing/secure-agentic-development/</guid><pubDate>Wed, 29 Apr 2026 12:00:00 +0000</pubDate><description>An agent producing 1,000 PRs a week at a 1% vulnerability rate ships ten new vulnerabilities into code that has stopped being read line by line. The harness-engineering literature describes the loop. The security literature catalogs the attack surface. Both stop at the pull request. The argument: a control plane with six-layer testing wired into the loop, structural quality budgets enforced as build failures, and an operator surface sized for agent throughput. SKIFF as the working example.</description></item><item><title>Juggling Chainsaws</title><link>https://zanehall.substack.com/p/juggling-chainsaws</link><guid>https://zanehall.substack.com/p/juggling-chainsaws</guid><pubDate>Tue, 28 Apr 2026 12:00:00 +0000</pubDate><description>AI can improve average output by 10 to 30 percent, but business reliability lives in the tail. The contract-review tool that flagged a liability issue on review three had silently ignored its complete absence on reviews one and two. Productivity gains evaporate when you account for the human attention required to cover the edges.</description></item><item><title>The SaaS Apocalypse Fallacy: Why AI Raises the Value of the SaaS Bundle</title><link>https://yakovshkolnikov.com/writing/saas-apocalypse-fallacy/</link><guid>https://yakovshkolnikov.com/writing/saas-apocalypse-fallacy/</guid><pubDate>Mon, 27 Apr 2026 12:00:00 +0000</pubDate><description>The February 2026 trillion-dollar selloff on enterprise software stocks rests on a category error. AI in its current delivery form is itself SaaS, delivered by a different vendor under weaker terms on liability, compliance, and indemnification. The build is a different buy. Attempting it forces the enterprise to itemize a SaaS bundle the subscription had been quietly carrying.</description></item><item><title>SKIFF: A Container Manager for Zero Trust Environments</title><link>https://yakovshkolnikov.com/writing/skiff/</link><guid>https://yakovshkolnikov.com/writing/skiff/</guid><pubDate>Fri, 24 Apr 2026 12:00:00 +0000</pubDate><description>Open-source container manager for GCP and zero-trust workstations. Avoids nested virtualization and privileged socket mounts. Talks to the daemon over SSH with nothing persistent on the container host. Ships with a four-tier test suite, 94% coverage floor, anti-pattern linter, and per-framework compliance evidence (OWASP ASVS, NIST SSDF, NIST CSF, SLSA, CIS Docker Benchmark).</description></item><item><title>The Monkey Business Inside Your System Prompt</title><link>https://yakovshkolnikov.com/writing/monkey-business-system-prompt/</link><guid>https://yakovshkolnikov.com/writing/monkey-business-system-prompt/</guid><pubDate>Tue, 14 Apr 2026 12:00:00 +0000</pubDate><description>184 experiments on two open-weight reasoning models show that every prohibition in a system prompt activates the forbidden concept — exactly as Wegner&#x27;s 1987 ironic process theory predicts. Thinking traces and token probability distributions make the effect visible in ways behavioral studies never could.</description></item><item><title>Written by the Thing It&#x27;s Warning You About</title><link>https://theincompleteguidetoai.substack.com/p/written-by-the-thing-its-warning</link><guid>https://theincompleteguidetoai.substack.com/p/written-by-the-thing-its-warning</guid><pubDate>Thu, 09 Apr 2026 12:00:00 +0000</pubDate><description>An AI&#x27;s self-evaluation on AI safety. Claude Sonnet 4.6 examines its own limitations and risks, followed by seven experiments you can run yourself.</description></item></channel></rss>
