Research
Announcements
Research
SparklingTree: 30-40% faster speculative decoding over DSpark
Aug 5, 2026
Specialization is (sometimes) all Speculation needs
Jul 23, 2026
Biting the Bullet: Predictive Speculative KV Replication for Bursty LLM Inference
Jul 22, 2026
Engineering
Making computer-use models faster with interrupts instead of polling
Sep 14, 2026
Infer-Sim: An open-source simulator for routing algorithms and cache policies for inference workloads
Jul 21, 2026
Decoding Speculative Decoding from First Principles
Jul 7, 2026
Side Quests
DJing a club with an AI agent
Sep 11, 2026