LLM Application Engineering

Design, build, evaluate and ship an LLM product end to end โ€” retrieval, tool use, agents, guardrails, evaluation and cost control. Ten weeks, part-time, mentor-reviewed.

โฑ 10 weeks๐ŸŽš Intermediate๐Ÿงฉ 3 projects + capstone๐Ÿ–ฅ Any device ยท our GPUs

๐ŸŽฏ Who it's for

Engineers and analysts who can already write Python and want to build real LLM applications โ€” not just call an API in a notebook. If you've finished Python & Maths for ML or have equivalent experience, you're ready.

โœ… Prerequisites

  • Comfortable writing and debugging Python (functions, classes, virtual environments in principle โ€” you won't set one up here).
  • Basic command of the terminal and Git.
  • Familiarity with HTTP APIs and JSON.
  • No prior machine-learning maths required for this course.

What you'll build

A RAG assistant

Retrieval-augmented assistant over a real document corpus, with its own retrieval evaluation set.

A tool-using agent

An agent that calls your code safely, with input validation, guardrails and spend limits.

A shipped capstone

A complete, documented LLM application on a real brief โ€” built, measured and presented.

Curriculum โ€” topics covered

Prompting & model I/O

  • System vs user prompts
  • Few-shot & pattern prompting
  • Reliable structured / JSON output
  • Streaming responses
  • Handling refusals & truncation
  • Token & context budgeting

Retrieval (RAG)

  • Chunking strategies
  • Embeddings & vector stores
  • Hybrid (keyword + vector) search
  • Re-ranking & metadata filtering
  • Citations & “I don't know” handling
  • Retrieval evaluation sets

Tool use & agents

  • Function calling & schemas
  • Input validation & safe execution
  • Planning loops & memory
  • Multi-step task decomposition
  • Agent failure modes & recovery
  • When not to use an agent

Evaluation

  • Building labelled test sets
  • Offline evaluation suites
  • LLM-as-judge and its pitfalls
  • Regression testing on prompts
  • Tracking quality over time
  • Reading eval results critically

Guardrails, safety & cost

  • Input / output filtering
  • Prompt-injection basics
  • Rate & spend controls
  • Response caching
  • Model selection for cost / latency
  • Graceful degradation

Production & iteration

  • Request tracing & observability
  • Collecting & using user feedback
  • Prompt & config versioning
  • Deployment of an LLM service
  • Finding the weak stage in a pipeline
  • Shipping an improvement with confidence

Week by week

Wk 1

Foundations sprint recap & the LLM app skeleton

Set up the workspace, call a model from code, structure a minimal app: prompt in, response out, logging around it.

Wk 2

Prompt design & structured output

System vs user prompts, few-shot patterns, getting reliable JSON, handling refusals and truncation.

Wk 3

Retrieval-augmented generation

Chunking, embeddings, a vector store, and wiring retrieval into the prompt. Project: a working RAG assistant over a supplied corpus.

Wk 4

Better retrieval

Hybrid search, re-ranking, metadata filtering, and building a labelled evaluation set for retrieval quality.

Wk 5

Tool use & function calling

Letting the model call your code safely: schemas, validation, error handling, and when not to.

Wk 6

Agents

Planning loops, memory, multi-step tasks, and the failure modes that make agents fragile. Project: a task agent with guardrails.

Wk 7

Evaluation

Offline eval suites, LLM-as-judge and its pitfalls, regression testing, and tracking quality over time.

Wk 8

Guardrails, safety & cost

Input/output filtering, prompt-injection basics, rate and spend controls, caching, and model selection for cost/latency.

Wk 9

Observability & iteration

Tracing requests, collecting feedback, finding the weak stage, and shipping an improvement with confidence.

Wk 10

Capstone

Ship an LLM application end to end on a real brief, present it, and hand over a documented, runnable project.

๐Ÿ Your capstone

“Build an internal assistant that answers policy and process questions for a mid-size company from its own documents. It must cite sources, say when it doesn't know, stay under a set cost per query, and come with an evaluation report showing accuracy on a 50-question test set.”

๐Ÿ“‹Example briefYou scope it, build it, measure it, and present it.

๐Ÿ’ณ Delivery & fees

Available in all four delivery formats โ€” self-paced, live online, hybrid weekend, and in-person. Fees vary by format; confirmed on your advisor call, instalments available. [Add concrete fees and the next cohort date here.]

Ready to build?

Apply for the next cohort, or book a call and we'll help you decide if this is the right starting point.

Apply now