What Actually Happens When You Send a Message to ChatGPT
From your keyboard to GPU clusters and back — the complete internal architecture of how LLMs process a single request at scale.
The Moment You Hit Send
You type a message. You press Enter. A response starts streaming back in seconds.
What just happened involves DNS routing, tokenization, GPU clusters, tensor parallelism, and a distributed system handling 900 million users. None of it is magic. All of it is engineering.
Here's the complete picture.
Stage 1 — The Request Leaves Your Device
Your prompt travels as an HTTPS request. Before it reaches any AI, it hits:
DNS Resolution → Your device resolves api.openai.com to the nearest IP via Anycast routing — a technique where one IP address maps to multiple physical servers globally. You're automatically routed to the closest edge location.
Global Router / Edge Layer → The request lands at an edge node. This handles TLS termination (decrypting HTTPS), DDoS protection, and initial load balancing. Latency here is measured in single-digit milliseconds.
Region Services → The request routes to a regional data center — the one closest to you with available capacity. Here sit the actual application servers: API gateways, authentication services, and rate limiters.
Key concept: Anycast + edge routing ensures no user is ever hitting a single data center. The network itself distributes load before any application code runs.
Stage 2 — Pre-Processing: Your Text Becomes Numbers
The model cannot read English. Before inference begins, your prompt goes through a pipeline:
Authentication & Rate Limiting → API key validated. Request quota checked. If you've exceeded your rate limit, the request is rejected here — before consuming any GPU resources.
Tokenization → Your text is broken into tokens. Not words — chunks. "Neural" might be one token. "Unbelievable" might be three. Each token maps to an integer ID.
"How does ChatGPT work?"
→ [2437, 1575, 12359, 38, 4, 6] (approximate token IDs)
GPT-4 uses roughly 100,000 token vocabulary. Every token is then converted into a high-dimensional vector — a list of floating-point numbers representing meaning in mathematical space.
Context Assembly → Your conversation history is retrieved from the session store (Redis or similar). Prior messages are prepended to your current prompt. The model sees the entire conversation, not just your latest message.
Key concept: The model has no persistent memory. Every request reconstructs the full conversation from stored session data.
Stage 3 — The Inference Scheduler: The Traffic Cop
Your vectorised prompt now enters the inference layer — the most complex part of the system.
The Problem: GPU clusters are expensive and finite. Thousands of requests arrive per second. You cannot send each request to a GPU individually — the overhead would be catastrophic.
The Inference Scheduler solves this. It acts as an intelligent dispatcher:
- Monitors load across GPU clusters
- Checks which clusters have relevant model weights already loaded in VRAM (warm vs cold)
- Groups incoming requests into batches
Continuous Batching is the key innovation here. Traditional batching waited for a full batch to assemble before processing. Continuous batching adds new requests to the GPU the exact millisecond a previous request finishes — keeping hardware at near 100% utilisation without increasing individual user latency.
This is the difference between a GPU running at 40% efficiency vs 95% efficiency. At scale, that gap is worth hundreds of millions of dollars.
Stage 4 — GPU Clusters: Where Inference Actually Happens
Your batch of requests reaches the GPU cluster. This is where the model lives.
The Memory Problem: GPT-4 has roughly 1.8 trillion parameters. Each parameter stored in 16-bit precision requires 2 bytes. That's ~3.6 terabytes — far exceeding a single GPU's VRAM (an H100 has 80GB).
Tensor Parallelism solves this. The model is split across multiple GPUs:
Layer 1-10 → GPU cluster A
Layer 11-20 → GPU cluster B
Layer 21-30 → GPU cluster C
GPUs communicate over NVLink (within a node) or InfiniBand (across nodes) — interconnects with bandwidth measured in terabytes per second. A single forward pass involves all these GPUs collaborating in real time.
The Forward Pass — what inference actually is:
Your token vectors pass through the transformer layers sequentially. At each layer, the model performs billions of matrix multiplications — calculating attention scores (which tokens should influence which other tokens) and transforming representations through feed-forward networks.
At the final layer, the model outputs a probability distribution over its entire vocabulary — 100,000 numbers, each representing the probability that word is the correct next token.
The highest probability token is selected. That's your first word.
Key concept: The model doesn't think. It performs a single, enormous mathematical function — input vectors in, probability distribution out.
Stage 5 — Autoregressive Generation: The Loop
Here's what makes LLMs fundamentally different from most software:
They generate one token at a time, and each new token becomes input for the next.
Step 1: "The" → model predicts "quick"
Step 2: "The quick" → model predicts "brown"
Step 3: "The quick brown" → model predicts "fox"
Every step requires a full forward pass through the entire model. A 200-token response means 200 complete forward passes.
KV Cache optimises this. During each forward pass, the model computes Key and Value matrices for the attention mechanism. Without caching, these would be recomputed from scratch on every step. KV Cache stores them — so each new step only computes attention for the newest token, not the entire history.
This is one of the most significant optimisations in production LLM serving. It reduces computation per step dramatically.
Stage 6 — Streaming the Response Back
Each generated token is immediately detokenized (converted back to text) and streamed to your client via SSE (Server-Sent Events) — a protocol where the server pushes data to the client over a persistent HTTP connection without the client polling.
This is why you see text appear word by word. The model isn't waiting to finish — it sends each token the moment it's generated.
Concurrency model: Each streaming response holds an open connection. At 900 million users, connection management becomes a distributed systems problem. Load balancers must maintain connection state. Servers must handle thousands of concurrent open streams without blocking.
Stage 7 — Quantization: Making This Affordable
Running FP16 (16-bit) inference at scale is extraordinarily expensive. Production systems use quantization — compressing model weights to lower precision:
- FP16 → INT8: 2x memory reduction, ~1% quality loss
- FP16 → INT4: 4x memory reduction, ~3% quality loss
A quantized model fits in less VRAM, runs faster, and costs significantly less per token. Most production inference runs on quantized models, not the full-precision research versions.
The Full Picture
You → DNS/Anycast → Edge → API Gateway
→ Tokenizer → Session Fetch → Scheduler
→ Continuous Batching → GPU Cluster (Tensor Parallel)
→ Forward Pass × N tokens (KV Cache)
→ Detokenizer → SSE Stream → You
Every component exists to solve a specific bottleneck. Anycast solves geographic latency. Continuous batching solves GPU utilisation. Tensor parallelism solves memory limits. KV Cache solves autoregressive overhead. Quantization solves cost.
There is no single clever idea here. It's a dozen engineering solutions layered on top of each other — each one essential at this scale.
That's what production AI engineering actually looks like.