Writing

Blog.

Engineering writing from the Veso team. Architecture, production playbooks, and field notes on shipping AI into legal, healthcare, and enterprise workflows.

Archive

All articles

2026
  1. AI Accelerator Options in 2026: NVIDIA, AMD, Google, AWS, Cerebras, Tenstorrent Twelve accelerators on memory per dollar, bandwidth per compute, MLPerf tokens per dollar, domain size, and software status.
  2. Kimi-Linear-48B on Tenstorrent Wormhole Veso AI open-sources the fused KDA, MoE, and MLA kernels that took Kimi-Linear-48B from not starting at all to a 16K-context server on four n300 cards.
  3. The Classifier That Watches Your Coding Agent Spawn a subagent in Claude Code and a second model reads the call before it runs, scores it 0 to 100, and blocks anything above 50. The full ruleset, captured off the wire and reproduced here in one piece.
  4. Knowledge as Dynamics: The Paper Veso R&D releases Knowledge as Dynamics: distilling, composing, and inheriting the latent vector fields of neural networks. Five laws, one teacher-only instrument, every number reproducible on a laptop.
  5. The LLM Timeline: What Lived and What Died Every pivotal LLM release from GPT-2 to Kimi K3, the six eras they define, and the techniques each era killed.
  6. Kimi K3: A Forensic Analysis A claim-by-claim audit of Kimi K3: architecture, benchmarks, pricing, and hardware.
  7. What We Learned Validating JEPA on a Laptop We reproduced the core claims of two 2025-26 JEPA papers (CrossJEPA and LeJEPA) at toy scale on an M5 MacBook Air. The mechanisms hold. Here is what that means for enterprises deciding whether self-supervised representation learning is real or hype.
  8. The Agentic Harness: Why the Orchestration Layer Is the Product Models are commodities. The harness, the control plane that governs what an LLM sees, calls, and outputs, separates demos from production AI systems. A technical breakdown of the architecture pattern defining 2026.