Research & Engineering Essays.
Rigorous empirical benchmarks, architectural teardowns, and analytical examinations of frontier AI models, autonomous agents, and inference systems.
On the Economics of Context: Latency, Cost, and Cache Invariance in Frontier LLMs
A systematic evaluation of KV-cache reuse on large prompt prefixes. Demonstrating how prompt structure and prefix invariance cut input token expenditure by up to 75% while reducing Time-To-First-Token from 1.4s down to 310ms.
The Fallacy of Heavy Agent Frameworks: Returning to Deterministic State Machines
Why multi-layered autonomous agent abstractions frequently fail in high-stakes production systems, and how minimalist typed state loops deliver superior reliability, observability, and deterministic bounds.
High-Dimensional Vector Search: Memory Geometry of HNSW vs Quantized Inverted Indices
A deep examination of graph-based versus inverted-file vector indexing when scaling beyond 1,000,000 dense vectors. Architectural trade-offs between DRAM footprint, re-indexing pauses, and NDCG recall.