Last month I described a research program organized around four areas. That structure still holds. But over the past few weeks a through-line became clear that I did not have language for in month one.
I am calling it architecturally assured AI.
The term comes from safety-critical engineering, where "high-assurance software" means correctness verified by design. For AI to drive real value it has to come through reliability and trustworthiness. And not at extraordinary cost. Assurance built into the architecture, not bolted on after the fact through audits and compliance theater.
Research
My first new preprint (with open source code) replaces three sources of training randomness with deterministic alternatives. Structured weight initialization instead of random sampling. Golden ratio batch scheduling instead of shuffling. Architecture choices that eliminate GPU non-determinism entirely. The result is bit-identical models across runs, with rare-class variance reduced by up to 7.5x.
The second preprint focuses on lower compute cost and more interpretable version of the attention for automatic speech recognition models. It does so by replacing self-attention with learned pulse gating functions for on-device speech recognition. I know people are patiently waiting for a Triton kernel. It is coming soon and with new training to get the error rate even lower.
Version 1.1 of the agent memory framework, persistent quantized KV caches for multi-agent inference on edge devices, is also in preparation. It is now targeting Jetson Thor as an additional platform to join MLX and in alpha release supports Claude Code CLI natively (and NemoClaw aspirationally).
I also demoed a classical robotics baseline for Franka Panda pick-and-place in Isaac Sim that show a classical low training cost solution to what many demos seem to be misusing for the baseline for RL.
And eleven new provisional patents this month, all built around architecturally assured AI for safety-critical and agentic systems. The first paper on this technology is in internal informal review right now.
Writing
I wrote about what the enterprise observability stack is missing for agentic AI. The foundational telemetry layer is taking shape.
Two new issues of The Incomplete Guide to AI on Substack. A guest article in Frictionless Data on why using the same AI to make the same decisions as your competitors converges everyone on the same strategy.
And plain-language explanations of both research papers on LinkedIn, because research that nobody outside the field can read changes nothing.
GTC and consulting
Almost a week at NVIDIA GTC. Picked up new consulting clients (some paid, some free just to give back to the world).
IRL: What's new since month one
- Two arXiv preprints (https://arxiv.org/abs/2603.28040, submitted to journal; https://arxiv.org/abs/2603.16922, Triton kernel coming soon)
- Eleven provisional patents
- Four LinkedIn articles (follow my profile)
- Two Substack issues (https://theincompleteguidetoai.substack.com)
- A guest article (https://zanehall.substack.com/p/automate-like-everyone-decide-like, Frictionless Data)
- One mentorship meeting
- Portfolio website launched (https://yakovshkolnikov.com)
