Students ask a course a question and get an answer that is actually from the course. Not a general-purpose model guessing at a syllabus it has never read — a retrieval pipeline that pulls the relevant passages out of the real material first, then asks the model to answer using only those.
How a question travels
A request enters through API Gateway and lands on a Lambda handler. The handler embeds the question, runs a semantic search over the vector index built from that course's documents, and passes the retrieved passages to the model as grounding context. Course documents live in S3; conversation state and per-course configuration live in DynamoDB.
Nothing is warm-started or pinned. Auto-scaling absorbed peak load — the spikes around assignment deadlines — with no manual intervention, which for a system with one maintainer was the difference between it working and it not.
Why retrieval and not fine-tuning
Course material changes every semester. Fine-tuning would have meant retraining against a moving target and re-validating every time an instructor swapped a reading. Retrieval keeps the model fixed and the corpus live: replacing a document is an S3 upload and a re-index, not a training run.
The tuning work went into the retrieval layer instead — how passages are chunked, how many are retrieved, how they are ordered in the prompt. Combined with prompt engineering on the grounding instructions, that is what brought the hallucination rate down.
Watching it run
A system serving that many people needs to be observable by the person on call for it, which was me. The admin dashboard reads DynamoDB streams and CloudWatch metrics to show usage in real time, query trends over the term, and — the part that turned out to matter most — knowledge gaps: clusters of questions the retrieval layer kept failing to answer well, which almost always pointed at material missing from the corpus rather than a bug in the pipeline.
Owning the whole thing
System design, deployment through CloudFormation, on-call, and the iteration loop after launch. The infrastructure is templated, so the environment is reproducible rather than something assembled by hand in a console and remembered.

