The Interview Edge Blog

Understand it.
Then explain it.

First-principles guides to the AI concepts that come up in serious technical interviews—built with diagrams, code, and numbers you can reason through.

9 deep dives · 6 company lenses

The study library

12 topics

01 / 12Core

RAG

How a model finds useful evidence before it answers—and where retrieval systems quietly fail.

ML systemsSystem design
→
02 / 12Core

Attention

The query–key–value mechanism, worked by hand before we touch a single line of PyTorch.

LLM internalsCore
→
03 / 12Advanced

RLHF

How human preferences become a training signal—from pairwise labels to a safer policy.

AlignmentML systems
→
04 / 12Advanced

KV Cache

Why autoregressive generation would be painfully wasteful without cached keys and values.

InferenceLLM internals
→
05 / 12Core

Evals

How to measure whether an LLM system is actually getting better—without gaming your test.

System designReliability
→
06 / 12Advanced

FlashAttention

Why attention is limited by memory movement, not math — and how tiling plus an online softmax fixes it.

LLM internalsSystems
→
07 / 12Core

Rate Limiters

How APIs say "slow down" without falling over — token buckets, sliding windows, and the distributed race.

System designML systems
→
08 / 12Advanced

Continuous Batching

How inference servers stop wasting GPU on padding — iteration-level scheduling and paged memory.

InferenceML systems
→
09 / 12Advanced

Model Degradation

How production models quietly rot — drift detection, golden evals, and the monitors that catch it.

ReliabilityML systems
→
10 / 12Core

URL Shortener

Shrink URLs without breaking the internet — base62 counters, the key service, and 91 terabytes of storage math.

System designBackend
→
11 / 12Core

Distributed Cache

The senior-to-staff question — consistent hashing rings, 3x replication, eviction, and taming the thundering herd.

System designBackend
→
12 / 13Core

YouTube

Two systems wearing one logo — async uploads, transcode farms, and why hot videos live on the edge.

System designVideo
→
13 / 14Core

Twitter Timeline

Push vs pull fan-out — precompute timelines for the many, merge celebrities in at read time.

System designBackend
→
14 / 14Core

Load Balancer

L4 vs L7, five algorithms, health checks, and the failure math — the traffic cop every backend interview expects.

System designBackend
→
15 / 15Core

Training-Data Pipeline

Crawl the web, keep the best tenth, kill duplicates, scrub PII — how LLMs actually get their data.

Machine learningData
→
No topics match that search. Try a broader concept or clear the filter.

Claude guides

from @theclaudecraft

LatestClaude Code

Anatomy of a .claude/ folder: 7 things inside

Every Claude Code project hides a .claude/ folder that quietly decides how Claude behaves: shared settings, private overrides, skills, subagents, slash commands, memory, and hooks — with the exact file templates.

→
SettingsClaude Code

7 hidden Claude Code settings that change everything

Buried in the settings files: the permission mode new sessions start in, hooks that guard your secrets, a custom status line, your default model, output styles, and auto-compact — with the exact JSON for each.

→
PluginsClaude Code

The code-modernization plugin, explained

Anthropic's official legacy-modernization plugin: an enforced assess → approve → build → prove sequence, five specialist agents, three build methods, and a byte-level proof that the new code behaves like the old.

→
PluginsClaude Code

The security-guidance plugin, explained

Anthropic's official security plugin: instant pattern warnings on every edit, an Opus 4.7 diff review each turn, and an agentic data-flow review on commit — with honest limits.

→
HarnessClaude Code

Skills vs subagents vs hooks

The three words everyone mixes up: reusable playbooks, fresh-context specialists, automatic tripwires — and which one you actually need.

→
PluginsClaude Code

5 official Claude plugins you didn't know existed

Anthropic's own marketplace plugins: deep security scans, codebase setup plans, CLAUDE.md care, the agent SDK kit, and real C++ code intelligence.

→
Case studyClaude Code

Spotify's shunt plugin, explained

How Spotify cut Claude Code token usage by a reported 90%: bulk reads and boilerplate rerouted to a cheap worker model — with honest limits.

→
ModelsClaude API

Opus 5.5 in plain English

Anthropic's new flagship: 20% cheaper tokens, 30%+ faster output, a 1M-token context window — and the benchmarks behind it.

→
HarnessAgents

The harness that won Anthropic's hackathon

How 286 skills and 68 subagents work as a pipeline — and the stealable patterns (review loops + memory vault) behind 260K GitHub stars.

→
AllLibrary

Browse the full Claude library

Every Claude guide in one place — new features, hidden GitHub tools, and agent harness patterns, explained simply.

→

Learn the mechanism,
not the buzzword.

Every guide starts with the smallest useful mental model, makes the arithmetic visible, then climbs toward production trade-offs. Interview prompts are framed as practice—not leaked question banks—so you learn to reason under follow-up, not memorize a script.