AI System Design
29 posts, 2024 to 2026
2026
- Parsing Documents: PDFs, Tables, Scans and Layout (opens on Medium)
- Cleaning and Deduplication Before Anything Is Embedded (opens on Medium)
- The Ingestion Pipeline: What You Ingest, What You Store, and What Starts the Rest (opens on Medium)
- A Baseline RAG and a Golden Set Before Any Optimisation (opens on Medium)
- The RAG Pipeline as Building Blocks: Seven Ways an Answer Goes Wrong (opens on Medium)
- A Framework for Deciding Where Retrieval Work Happens (opens on Medium)
- Caching for LLM Systems: Exact, Semantic, and Provider Prefix Caches (opens on Medium)
- Reliability & Fault Tolerance in LLM Systems: Fallbacks & Guardrails (opens on Medium)
- Reliability for Model-Locked Systems: When Cross-Model Fallback Isn't an Option (opens on Medium)
- Evaluating LLM Systems: Offline Evals, Online Evals, LLM-as-Judge (opens on Medium)
- LLM Serving Architectures: Batching, KV-Cache, Multi-Tenancy (opens on Medium)
- Orchestration & Memory: Context Management for Long-Running Agents (opens on Medium)
- Agentic Systems: Tool Use, Planning, Multi-Step Reasoning (opens on Medium)
- Fine-Tuning vs. RAG vs. Prompting: Choosing the Right Approach (opens on Medium)
- Designing RAG Systems: Retrieval, Chunking, Re-ranking, Grounding (opens on Medium)
- Prompt Engineering as a System Design Discipline (opens on Medium)
- Why Framing Comes Before Architecture (opens on Medium)
- Why LLM-Era AI Systems Break Every Rule You Learned About ML in Production (opens on Medium)
- Beyond Accuracy: A Developer's Guide to Reliable LLM Evaluation (opens on Medium)
- How to Stop Your NL2SQL Agents From Crashing in Production: The Worker-Pool Pattern (opens on Medium)
- Beyond Schema: Why Your AI Can't Write Good SQL (and How to Fix It) (opens on Medium)
- Engineering Trust: A Defensive Architecture for NL2SQL Systems (opens on Medium)
2025
- Scalable Inference with RDMA and Tiered KV Caching (opens on Medium)
- The 70B LLM Optimisation Playbook: From 57.5GB to 24.3GB Per GPU (opens on Medium)
- How to Serve a 70B Model with a 128K Context on Just 8 H100s (opens on Medium)
- Decoding Real-Time LLM Inference: A Guide to the Latency vs. Throughput Bottleneck (opens on Medium)
2024
- Discover the Magic Behind Your Searches: How Semantic and Vector Search Transform Your Online Experience (opens on Medium)
- The Mechanics of Query Expansion in RAG Systems: A Theoretical Exploration of PRF and LLM Techniques (opens on Medium)
- Rediscovering Query Expansion: The Classic Technique Powering Modern AI Searches (opens on Medium)