The Lab
Soundness AI is my independent research lab. Its premise: AI does not need to be infallible to be useful, but the systems around it must stay trustworthy when it fails. This page collects the progress updates and shows how to do AI research without a corporate budget: the current setup and what each piece is for.
Updates
Seven Months In: Soundness AI Update
Seven months of Soundness AI. Putting a human in the loop is not enough to earn the trust critical work requires. 7 research and working papers, 14 provisional applications across 21 invention families, 30+ articles, and a large integrated trust prototype in development.
Two Months In: Independent Research Lab Update
The research program found its through-line: architecturally assured AI. Reliability as a structural property, not a compliance exercise.
One Month In: IRL (Independent Research Lab) Update
What an independent AI research lab actually looks like. Hardware, budget, what worked, what didn't.
The Essential Beauty of Abstracting Physical Experiments
Mashmeasure, the system I built in the Princeton lab in 2003. Experiments written in natural units while auto-discovered instrument shims enforced limits, rates, and safe operation. The deterministic execution layer agents still need.
Patents
- 14 Provisional Patent Applications, Soundness AI
- 14 provisional patent applications covering 21 invention families (2026), currently unassigned.
For commercial products or licensing, contact me on LinkedIn.
Setup
Constrained Compute Forces Better Research. A 10GB GPU and a $200/mo Claude subscription is not a limitation. It is a forcing function. If your experiment needs an H100 cluster to show results, the result probably isn't robust. Every system I build is designed to run on hardware a single researcher can afford.
Local compute
- MacBook Pro M5 Max 18-Core CPU / 40-Core GPU / 128GB
- 614 GB/s unified memory bandwidth. Metal acceleration for MLX. Large models and trustworthy AI work.
- MacBook Pro M4 Pro 12-Core / 24GB
- 273 GB/s unified memory bandwidth. Metal acceleration for MLX. Runs 7B-12B quantized models locally for inference and LoRA fine-tuning. Paged attention, batch inference, privacy-sensitive medical AI work.
- NVIDIA Jetson Thor
- Blackwell-based edge AI module. 2070 TFLOPS FP4 sparse, 128GB LPDDR5X at 273 GB/s, 14-core Arm Neoverse V3AE. 40–130W configurable. Local LLM inference for on-device agentic workloads.
- RTX 3080 (10GB VRAM), CUDA Workstation, 64GB RAM
- Mostly edge and robotics work. CUDA training for sub-3B models. QLoRA fine-tuning. Isaac Sim for robotics simulation (Franka pick-and-place experiments). The 10GB ceiling forces discipline about what actually needs GPU training vs. what doesn't.
Cloud GPU
- Lambda Labs, A100 Cloud
- Secure and confidential GPU rental up through GH200. On-demand access for experiments that exceed local VRAM. Pre-configured PyTorch environment, SSH in and run.
- Vast.ai, Budget GPU Rentals
- Marketplace for renting consumer and datacenter GPUs at a fraction of cloud pricing. Good for longer training runs where you need hours, not minutes. Setup guide coming soon.
Models and APIs
- Claude Code (Max Plan, $200/mo)
- Primary development environment. Agentic coding, long-duration workflows, orchestration framework development. The tool I use to build tools.
- Anthropic Claude API
- Production-grade inference for agentic workflows and orchestration framework. Sonnet for high-volume, Opus for complex reasoning chains.
- OpenAI API
- Used alongside Claude. Multi-vendor orchestration testing. The agentic framework was designed to be model-agnostic, so testing across providers is part of the methodology.
- Google Gemini API
- Used alongside Claude and OpenAI for multi-vendor orchestration testing and cross-provider comparison.
- DeepSeek API
- Low-cost inference for high-volume experiments and benchmarking. Useful for running comparison baselines against local models.
Frameworks and simulation
- MLX (Apple Silicon ML Framework)
- Apple's native ML framework optimized for Metal. Used for local inference, LoRA fine-tuning on M4 Pro, and running quantized models for privacy-sensitive medical AI experiments without cloud dependency.
- PyTorch + CUDA
- The standard. When a paper publishes results, this is how you verify them. Also for writing quick custom kernels when a customer problem needs something that doesn't exist in a library yet.
- NVIDIA Isaac Sim
- Robotics simulation for classical-vs-learned baseline experiments. Franka Panda pick-and-place, sensor simulation, reinforcement learning environments. Runs on the RTX 3080.



