Skip to content
learnwithF.I.R.E
0%
Booting FIRE
CoreIntermediate

Agentic AI Engineering

A 12-week program for engineers who already write Python and want to build AI agents that survive production. You'll build a multi-agent system, an evaluation harness that catches its failures, and a deployment with full tracing — then defend all three in a technical review.

Duration
12 weeks
Level
Intermediate
Format
Live online
Language
English
Investment
Request pricing
Next cohort
Dates announced soon
Credential
FIRE Certified Agentic Engineer

Who this is for

This is not for

  • Complete beginners — start with Python for AI Engineering instead
  • Anyone looking for a no-code or prompt-only course
  • Anyone who wants a certificate without shipping four working artifacts
  • Teams needing a vendor-specific certification — we teach the stack, not one API

What you leave with

Tools and stack

Python 3.12LangGraphAnthropic and OpenAI APIsOpen-weight models via Ollama and vLLMPydantic for structured outputFastAPIPostgreSQL with pgvectorOpenTelemetry and LangfuseDockerpytest

Curriculum

Module 1 — What an agent actually is

  • Define an agent precisely: a loop with tools, state, and a stopping condition
  • Identify the three cases where a plain LLM call beats an agent
  • Build a minimal agent loop from scratch, without a framework
  • Explain why most "agent" demos fail the moment inputs get messy

Module 2 — Reliable model output

  • Force structured output with Pydantic schemas and validate every response
  • Handle refusals, truncation, and malformed JSON as expected states, not bugs
  • Measure output variance across runs and decide what variance you can tolerate
  • Choose temperature and model tier from evidence rather than habit

Module 3 — Tool design and function calling

  • Design tool interfaces an LLM can actually use correctly
  • Write tool descriptions that reduce wrong-tool selection
  • Handle tool errors, timeouts, and partial failures inside the loop
  • Apply least privilege: give each tool the narrowest scope that works

Module 4 — Memory and state

  • Separate working state, conversation history, and durable memory
  • Manage the context window deliberately: what to keep, summarise, or drop
  • Choose between structured storage and vector storage per data type
  • Diagnose failures caused by state, not by the model

Module 5 — Retrieval for agents

  • Build a retrieval pipeline: chunking, embedding, hybrid search, reranking
  • Let the agent decide when to retrieve instead of retrieving on every turn
  • Measure retrieval quality separately from answer quality
  • Recognise the three failure modes RAG cannot fix

Module 6 — Multi-agent orchestration

  • Implement planner/worker and supervisor patterns in LangGraph
  • Design clean handoffs and shared state between agents
  • Justify when one agent beats several — most systems are over-decomposed
  • Control runaway loops with budgets, depth limits, and termination conditions

Module 7 — Evaluation harnesses

  • Build a golden dataset from real failures, not invented examples
  • Write deterministic assertions before reaching for LLM-as-judge
  • Understand where LLM-as-judge is unreliable and how to calibrate it
  • Run evals in CI so a prompt change cannot ship a regression

Module 8 — Guardrails and human-in-the-loop

  • Place approval gates where the cost of a wrong action is irreversible
  • Set confidence thresholds and design the escalation path
  • Build an audit trail that reconstructs why the agent did what it did
  • Distinguish guardrails that add safety from those that only add latency

Module 9 — Observability

  • Instrument the agent with OpenTelemetry spans across the full loop
  • Track tokens, latency, and cost per run and per user
  • Build a failure taxonomy from production traces
  • Set alerts on the signals that predict user-visible failure

Module 10 — Security and red-teaming

  • Execute prompt injection attacks against your own agent, direct and indirect
  • Defend tool access: allowlists, scoping, confirmation for side effects
  • Treat all retrieved and web content as untrusted input
  • Write a threat model for an agent with real system access

Module 11 — Deployment

  • Run in shadow mode against live traffic without acting on it
  • Ship behind a canary with an explicit rollback trigger
  • Control cost with caching, model routing, and hard budget limits
  • Decide what runs on a frontier API and what runs on an open-weight model you host

Module 12 — Capstone build and technical review

  • Ship your agent, harness, and observability as one working system
  • Present the architecture and defend the trade-offs to a technical panel
  • Analyse your own failure data and propose the next iteration
  • Leave with a portfolio artifact you can walk an interviewer through

How this program handles evaluation and governance

Responsible engineering isn't a module in this course — it's four of them.

Week 7 makes evaluation a build requirement: no agent ships without a golden dataset and assertions running in CI. Week 8 covers approval gates, confidence thresholds, and audit trails detailed enough to reconstruct any decision the agent made. Week 10 has you attack your own system with prompt injection and tool abuse before anyone else does. Week 11 covers shadow-mode deployment — running against live traffic without acting on it, so failure costs you nothing.

The reason this is spread across the course rather than bolted on at the end: governance added after a system is built is documentation. Governance designed into the loop is engineering.

Questions people ask

Do I need machine learning or maths background?

No. This is a systems engineering course, not an ML theory course. You need to write Python confidently and reason about distributed systems. We don't derive gradients or train models from scratch.

How much time per week?

Around 8 hours: one live session, one working session, and independent build time. The capstone weeks run heavier. If you can't protect that time, take a later cohort rather than falling behind in this one.

Which models and frameworks do you teach?

LangGraph for orchestration, Anthropic and OpenAI APIs, and open-weight models run locally with Ollama and vLLM. We teach the patterns deliberately across more than one provider, because the APIs change and the patterns don't.

Is this live or recorded?

Live online, with sessions recorded for review. Code review and the capstone panel are live and interactive — that's where most of the learning happens.

What do I have at the end?

Four artifacts: a multi-agent system, an evaluation harness, a deployed agent with tracing, and a written failure analysis. All of it yours, on your GitHub, and specific enough to discuss in a technical interview.

Do you guarantee placement?

No, and be sceptical of anyone who does. We provide portfolio review, interview preparation, and introductions where there's a genuine fit. The artifacts do the work — that's why the capstone is assessed by a panel rather than auto-graded.

Can my employer sponsor this?

Yes. We'll provide a scope document and invoice for reimbursement. For three or more people from one organisation, look at the enterprise track instead — it's delivered against your own systems and use cases.

What if I fall behind?

Session recordings and office hours cover a missed week. Miss more than three and we'll move you to the next cohort at no charge — finishing badly serves nobody.

Related courses

Cohort forming — join the waitlist

No payment at this stage. We confirm dates, fees and fit before you commit to anything.

Join the waitlist
Register Now