Inside Hamba's agent-run newsroom: choosing models for each role, tracing every run, and why human taste remains the most important part of the pipeline.
In 2008 I built a travel website called Hamba, in PHP, the way we built everything back then. It lived, it aged, it was forgotten.
Years later I was cleaning up old disks and found the codebase. For fun, I had Copilot port it to .NET, and here's the strange part: it worked. The whole site, alive again in an afternoon. But looking at it, I realized something more interesting: the code was fine, the thinking was outdated. It was a website built for a world where humans wrote every page by hand. So instead of shipping the resurrected version, I asked a better question: if I built Hamba today, from zero, in the age of agents, what would it be?
The answer is hamba.nl, a Dutch travel magazine where the editorial pipeline is run by a team of AI agents. It's a real production system, not a slide demo. And running it for a while has taught me more about working with agents than any tutorial, so let me give you the tour.
The newsroom
Think of it as a small newsroom where the staff happens to be digital, and every employee has exactly one job.
There are agents that do the thinking: research, planning what a piece needs, deciding structure. There are writer agents, one for Dutch and one for English, and I'll be honest: the prompt engineering for those two took more work than anything else in the system. Getting an agent to write is easy. Getting it to write like your magazine, in two languages, without slipping into the generic AI voice, is a craft. There's an agent that makes the images. And there's a validator that checks every piece against the editorial rules before anything moves forward, because an editorial pipeline without a quality gate is just a content hose.
At the end of the pipeline sits the most important part: a human. Nothing publishes without approval. More on that in a minute, because it's not the disclaimer it sounds like. It's the design.
Hire the model for the role
Here's something running this taught me that demos never show: you don't pick one model, you hire a team.
The model that's great at thinking through what an article needs is not the model I want writing the prose. The one that writes beautiful Dutch is not automatically the one for images. So each role in the newsroom runs the model that's best at that job, and I test candidates the way you'd interview new hires: give them the actual work, compare the results, pick per role. Which models? That answer changes every few months, and that's exactly the point.
Build the newsroom so you can swap an employee without rebuilding the office.
If that sounds like my hiring advice for humans, that's not an accident. One job per agent, the right hire per role, and judgment stays with whoever has actually done the work.
What agents are like as colleagues
After months of running this team, here's my honest performance review.
Great at: volume, consistency, and never having a bad Tuesday. The pipeline produces at a pace no human newsroom this size could match, in two languages, every piece following the same editorial rules.
Terrible at: taste. An agent doesn't know that a piece is technically fine but boring. It doesn't know that this opening has been used a hundred times, or that a place description sounds like it was written by someone who has never been anywhere. That judgment, the difference between correct and good, still lives entirely on the human side of the pipeline.
That's why the approval step is the design and not a disclaimer. The agents take the volume. The human keeps the taste. Sound familiar? It's the same line I draw everywhere: automate the task, never the point. The point of a magazine is that it's worth reading, and worth reading is a human judgment.
You can't manage a team you can't see
The least glamorous lesson, and maybe the most important one for anyone building multi-agent systems: observability is not optional.
When a piece comes out wrong, I need to know which agent did what, in what order, with what input. So every run is traced, and I built a visualizer that can replay a full pipeline run step by step. It's the agent version of an editorial meeting: what happened, where did it go sideways, who needs better instructions.
If you can't see what your agents did, you don't have a team. You have a mystery that occasionally publishes.
Why bother?
A fair question: why run a real magazine instead of just building demos?
Because production teaches what demos can't. A demo has to work once, on stage, for ten minutes. A magazine has to work every day, with real content, real edge cases, and real consequences when the Dutch writer agent suddenly develops a strange new habit. Every hard lesson in this post came from the gap between those two. I've said before that you should build the thing before you demo the thing. hamba is me doing that with the biggest thing I talk about on stage: a workforce of agents.
And the meta-lesson, the one that keeps coming back in everything I write: the more the agents produce, the more valuable the human at the end becomes. I didn't automate myself out of a magazine. I automated myself into the only job that matters there: deciding what's good.
The 2008 version of Hamba needed me for everything, so it died when I got busy. The 2026 version needs me for one thing, so it lives.
Behind the newsroom
Explore the slides
Your new teammates. Architecting a multi-agent workforce.
A closer look at Hamba's agents, architecture, and editorial workflow. Browse the slides here in the article.
Read slide text
Your new teammates.
Architecting a multi-agent workforce.
Henk Boelman
Principal Developer Advocate, Microsoft
Read slide text
Hi, I'm Henk.
Principal Developer Advocate at Microsoft
henkboelman.com
in/henkboelman
hnky
youtube.com/henkboelman
@hboelman
henkboelman.com
Read slide text
The story about Hamba
hamba · isiZulu: to go, to travel. hamba kahle · go well.
henkboelman.com
Read slide text
Hamba, since 2008
web.archive.org · www.hamba.nl · 29 Oct 2010
2008
Built it.
A travel platform, hand-written in PHP. Explore, travel, Hamba TV, a shop, My Hamba.
LATER
Found it.
On an old disk, while cleaning up. Not responsive, not maintained, not proud.
COPILOT
Ported it.
Handed the repo to Copilot: port this to .NET. It did, and it worked.
BUT
Still 2008.
Same pages, same ideas. A faithful port of something that no longer made sense.
2026
Rethought it.
Not a port this time. A platform with a memory, built for readers and for agents.
Copilot can port anything. It cannot tell you that you should not.
henkboelman.com
Read slide text
Hamba
hamba.nl/bestemmingen/afrika
WHAT
Dream. Explore. Discover. Share.
A bilingual travel magazine at hamba.nl: regions, countries, hotspots and stories, with real readers.
WHY
A platform with a memory.
Travel facts and travellers' stories usually live apart. Hamba puts them around the same destination.
WHO
Written by agents.
Nine of them: brief, research, two writers, review, links, hotspots, images. Travellers keep their own voice. A human approves.
Hamba means travel in Zulu. Travel and stay well.
henkboelman.com
Read slide text
How a story gets made
THE WORKFORCE
01
Plan
Turn one line into a brief: the angle, what must be covered, what to find out.
02
Research
Answer every open question, with sources it can point at.
03
Write
Two languages at once, from the same brief. Neither draft sees the other.
04
Edit
Review, validate, resolve every place name, make the images.
05
Publish
Save the story, attach the hero, close the assignment.
What stays human
IN
The assignment.
One line from an editor, in either language.
OUT
The approval.
One click, before a single reader sees it.
Everything in between is the workforce.
henkboelman.com
Read slide text
Meet the team
H
Henk
EDITORIAL · PUBLISHER
EDITOR IN CHIEF
Keeps the course, checks every contribution and decides what Hamba publishes.
M
Mara
EAST AFRICA · NATURE
AI AGENT
Follows a landscape past the well-known hotspot and works out how places connect.
J
Joris
EUROPE · RAIL
AI AGENT
A soft spot for night trains, ferries and routes where getting there is half the trip.
A
Aisha
NORTH AFRICA · FOOD
AI AGENT
Starts at the market and uses food to find stories about neighbourhoods, traditions and people.
T
Tomoko
EAST ASIA · CITIES
AI AGENT
Dives into metro maps, changing neighbourhoods and smart ways to live a big city locally.
S
Sofia
LATIN AMERICA · CULTURE
AI AGENT
Looks for the culture behind the cliché and ties what happens today to a place's history.
D
Daan
OUTDOORS · ROUTES
AI AGENT
Thinks in trails, seasons and detours, preferably off the busiest route.
E
Elias
CLIMATE · TRAVEL FACTS
AI AGENT
Loves the practical layer: weather, travel time, transport and the details that make a plan usable.
One writer agent, seven voices. personality_prompt rides with the assignment into the writer prompts. The brief-maker and researcher never see it.
The staff you would otherwise have hired. Eight bylines on hamba.nl/over/redactie, seven of them personas.
henkboelman.com
Read slide text
The rebuild
THE SITE
hamba.nl
.NET · SQL · containers on Azure
Readers, travellers and the editor. Nothing here knows an agent exists.
THE DOOR
Hamba MCP
sixteen tools, part of the site
The only way in for software. Read the catalogue, take an assignment, save a story.
THE TEAM
Nine agents
Agent Framework · Foundry Hosted Agents
Brief, research, write, review, link, illustrate. Each one its own container.
MCP · streamable HTTP
A2A · agent cards, tasks
HOW
01
Described, not typed.
The site, the MCP server and the agents were built in GitHub Copilot from a description of what they should do.
02
Rules first.
Before the first feature: an AGENTS.md, the skills, the specialist agents. Teach the team, then let it build.
03
Shipped as containers.
The site in its own containers, every agent in its own, all on Azure, all deployed the same way.
Next: the same method on a small demo blog, so the files fit on a screen. The real repo is bigger, the rules are the same.
Same name, same domain. Nothing else survived.
henkboelman.com
Read slide text
Demo
Hamba
hamba.nl · Dutch and English · nine bylines, one editor
1 OF 4
henkboelman.com
Read slide text
DEMO 1 OF 4 · RECAP
Hamba
01
It publishes.
Real articles, real readers, a real domain. Not a sandbox that resets when the talk ends.
02
Two languages, not a translation.
One brief, two writer calls, two prompts. Neither draft ever sees the other one.
03
Every place is a link.
No writer typed those. The linker resolves place names against the catalogue, and writes the page when it is missing.
04
The masthead is honest.
Eight bylines, seven of them personas, and the site says so in plain Dutch.
That is the product. The rest of the hour is the workforce behind it.
henkboelman.com
Read slide text
The rest of the hour
01
Building with GitHub Copilot
How do you teach a tool the rules of your house?
DEMO 2
02
Making Hamba ready for Agents
How does a website open a door that software can walk through?
03
Building the agents
What is an agent actually made of, once you take the hype out?
04
Agents talking to agents
How do nine of them stay in step without a group chat?
DEMO 3
05
Running it
Where does it live, and how do you see what it did at three in the morning?
DEMO 4
Four demos. One of them already happened.
henkboelman.com
Read slide text
Building with GitHub Copilot
henkboelman.com
Read slide text
Choose your IDE
Same repo. Same skills. Same agents.
henkboelman.com
GitHub Copilot App
AGENTIC
Hand it an issue, get back a branch. It plans, edits across files and opens the pull request.
VS Code
INTERACTIVE
Chat and edits in the file you are already in. Where you take the wheel.
GitHub CLI
SCRIPTABLE
The same agents from a terminal or a workflow, with nobody in the chair.
Read slide text
How the team is taught
INSTRUCTIONS
The house rules.
Always loaded. Short. Applies to everything.
“We write tests. We use xUnit. Dutch content is written in je/jij form.”
SKILLS
The procedures.
Loaded when relevant. Detailed. Any agent can use them.
“Here is exactly how you add a tool to our MCP server.”
AGENTS
The specialists.
Called by name. Own persona, own tool boundary. Reasons and plans.
“@pipeline-debugger, run 8f3a failed after the translator step.”
henkboelman.com
Read slide text
Demo
Teaching a repo to behave
an empty repo · one house rule · four minutes
2 OF 4
henkboelman.com
Read slide text
Side by side
Instructions
Skills
Agents
WHAT
Always-on rules
Repeatable how-to
Specialist role
LOADED
Every request
When relevant
When you call it
FILE
AGENTS.md
.github/skills/*/SKILL.md
.github/agents/*.agent.md
TEST
COVERAGE
“We aim for high coverage on src/”
How to write, run and report tests here
@test-engineer closes the real gaps
Rules.
Recipes.
Roles.
henkboelman.com
Read slide text
Making Hamba ready for Agents
henkboelman.com
Read slide text
What is MCP
An open protocol between agents and their tools and data. A server describes what it offers, any agent can discover and call it.
WITHOUT MCP
Every agent wired to every tool.
Writer
Fact-checker
Moderator
Storage
Images
Search
3 × 3 = 9 integrations. Add a tool, touch every agent.
WITH MCP
One protocol in the middle.
Writer
Fact-checker
Moderator
MCP
Storage server
Images server
Search server
3 + 3 = 6. Add a server, every agent gets it.
A server exposes tools it can run, resources it can read, and prompts it recommends. Typed, discoverable, the same on every surface.
henkboelman.com
Read slide text
MCP went stateless.
Any request, any instance.
No handshake. No session id. Nothing on the connection to lose.
henkboelman.com
Read slide text
Old vs new spec
2025-11-25
Stateful. Sessions, handshakes, callbacks.
2026-07-28
Stateless. One request at a time, any instance.
SESSIONS
initialize handshake, then Mcp-Session-Id on every call. State lives on the connection.
Gone. Protocol version and capabilities travel in _meta on each request.
DISCOVERY
Capabilities exchanged once, during the handshake.
server/discover: one call for versions, capabilities, identity. Before anything, or never.
SERVER NEEDS INPUT
Server calls the client back over an open stream: sampling, elicitation, roots/list.
Multi round-trip. Result says input_required, client retries the same request with answers.
NOTIFICATIONS
HTTP GET stream plus resources/subscribe. Resumable with Last-Event-ID.
subscriptions/listen: one opt-in stream, tagged by type. Broken stream, re-issue the request.
LONG-RUNNING WORK
Experimental tasks inside the core. Blocking tasks/result.
Tasks as an extension. Poll tasks/get, feed tasks/update. Servers can hand back a task unasked.
Deprecated: roots, sampling, logging, HTTP+SSE, dynamic client registration. Twelve months to move.
henkboelman.com
Read slide text
Agent to MCP
Agent
MCP server
server/discover What can you do?
versions, capabilities
tools/list What can I call?
typed schemas
tools/call Do this.
result
Every request carries its protocol version and capabilities in _meta. No session to lose between arrows.
henkboelman.com
Read slide text
Hamba MCP
ASSIGNMENTS
the queue
take_story_assignment
get_story_assignment
complete_story_assignment
whoami
CALLED BY
travel-writer
CATALOGUE
read only
list_regions
list_countries
list_hotspots
get_country
get_hotspot
CALLED BY
destination-linker · hotspot-writer · country-writer
CONTENT
writes
save_story
save_hotspot
save_country
update_preview
CALLED BY
travel-writer · hotspot-writer · country-writer
MEDIA
the library
import_media
import_story_image
import_country_image
CALLED BY
travel-writer · hotspot-writer · country-writer
allowed_tools per agent. The destination-linker can read the catalogue and nothing else. Nobody talks to the database.
The website is the MCP server. Sixteen tools, and the agents are its only writers.
henkboelman.com
Read slide text
Demo
Adding MCP
an empty repo · one house rule · four minutes
henkboelman.com
Read slide text
Building the agents
henkboelman.com
Read slide text
The harness
THE HARNESS
what you build
instructions
who it is, how it writes
tools
only the ones it may touch
middleware
validate, retry, refuse
telemetry
spans, payloads, a trace id
THE MODEL
one slot
AZURE_AI_MODEL_DEPLOYMENT_NAME
loop
The model is one slot.
Terra to think, Sol to write, Image-2 for pictures. Changing it is an environment variable and a re-evaluation, not a rewrite.
The loop is the agent.
Call, tool calls, results, call again, stop. The interesting question is who decides to stop. In hamba the workflow decides, never the model.
Everything else is yours.
Instructions, scoped tools, middleware that validates and retries once, spans that record what actually crossed the wire.
Everyone asks which model. The answer that matters is what you put around it.
henkboelman.com
Read slide text
What to choose
FRAMEWORK
What runs the team?
Considered An LLM coordinator agent, Semantic Kernel, AutoGen, hand-rolled orchestration.
Chosen Agent Framework, deterministic workflow
One @workflow function in plain Python controls every transition. Each @step is checkpointed. Collaborators over A2A, the website over MCP. No model ever owns the flow.
MODELS
One model, or one per job?
Considered One frontier model for everything, or a different one per role.
Chosen GPT-5.6 Terra to think, Sol to write
Briefing, research, review and linking on Terra. English and Dutch drafts on Sol, separate prompts, never translated. Images on GPT-Image-2. All of it in env config.
WHERE TO HOST
Where do the agents live?
Considered Inside the website, Container Apps jobs, Foundry Hosted Agents.
Chosen Nine Foundry Hosted Agents, azd deploy each
Every agent is a container with a Responses and an A2A endpoint, an agent card, and a managed identity. The website stays a plain .NET app with an MCP server.
HOW TO MONITOR
How do you know it worked?
Considered Container logs, Foundry Observability, Application Insights.
Chosen Application Insights, plus a trace visualizer
Every A2A call and MCP write is a span. Request and response payloads are recorded exactly, chunked and hashed. Paste one operation id, replay the whole editorial run.
henkboelman.com
Read slide text
Agent design
agents/researcher · one agent, six parts
ROLE
prompt.md
Answers one standalone question with two-phase web research. Never plans the article.
MODEL
RESEARCHER_MODEL=gpt-5.6-terra
An environment variable. The prompt does not know which model reads it.
IN
ResearchRequest
One question, the assignment, the content purpose. Never other questions, prior answers or the writer's identity.
OUT
ResearchQuestion
Answer, status, sources. Source ids are validated in code before anything downstream trusts them.
TOOLS
Web IQ MCP, budgeted
Ten searches, ten browses, two seconds apart, twenty-second timeout. Enforced by middleware, not by asking nicely.
CARD
protocols: responses, a2a
Discoverable by name from the project. The orchestrator finds it, it never finds the orchestrator.
01
One job, one agent.
The researcher answers one question. The brief-maker never sees the writer's personality. Narrow inputs are the design, not a limitation.
02
Contracts, not prose.
Pydantic in, Pydantic out, response_format on every model call. The orchestrator validates. No model parses another model's text.
03
Tools scoped per agent.
destination-linker gets a read-only Hamba catalogue and can wrap text in links but cannot rewrite it. Budgets live in middleware.
04
Preview by default.
hotspot-writer and country-writer return a candidate first and write only on save=true. Nothing persists without an explicit flag.
henkboelman.com
Read slide text
Shapes of a team
Sequential
SequentialBuilder
One after another, fixed order. Each agent sees what the one before it produced.
Concurrent
ConcurrentBuilder
Same input to all of them, results merged at the end. Independent work only.
Handoff
HandoffBuilder
One agent holds the work and passes it on when it is not the right one. Triage, then a specialist.
Group chat
GroupChatBuilder
Everyone in one conversation, a manager picks who speaks next. Critique, debate, consensus.
Magentic
MagenticBuilder
A manager writes a plan, delegates, tracks what came back and replans. Open-ended problems.
Your own workflow
@workflow @step
You write the control flow. Every transition is Python you can read, checkpoint and replay.
HAMBA
Five come with the framework. The sixth you write yourself. The more the manager decides, the less you can replay.
henkboelman.com
Read slide text
How they work together
Hamba MCP
take_story_assignment
hamba.nl/mcp
travel-writer
one @workflow, checkpointed @steps
no model
brief-maker
writing instruction + research plan
Terra
researcher ×n
one question each, Web IQ
Terra
writer
English ∥ Dutch, two prompts
Sol
editorial-reviewer
finalize each language, verdict
Terra
validate
slugs, sources, links, length
one correction retry
destination-linker
read-only catalogue, wrap links
Terra
hotspot-writer
missing hotspots, parent first
Terra
image-creator
hero + inline, private blob
GPT-Image-2
Hamba MCP
import media library
hamba.nl/mcp
Hamba MCP
save draft, complete assignment
hamba.nl/mcp
Hamba website over MCP, direct call_tool
collaborator agent over A2A, typed contract
workflow code, no model
henkboelman.com
Read slide text
Microsoft Agent Framework
HOSTING
Runs it as a Foundry hosted agent.
IN THE FRAMEWORK
ResponsesHostServer · agent_framework_foundry_hosting
IN HAMBA
Nine containers. Each main.py wraps one Agent or one workflow in a host that speaks Responses and A2A.
WORKFLOW
Decides the order. Code, not a model.
IN THE FRAMEWORK
@workflow · @step · workflow.as_agent()
IN HAMBA
travel-writer: one async function, thirteen checkpointed steps, parallel EN and NL inside one step.
AGENT
One role, one contract, scoped tools.
IN THE FRAMEWORK
Agent(client, instructions, tools, middleware, default_options)
IN HAMBA
brief-maker, researcher, writer, reviewer, linker: a prompt.md, a Pydantic response_format, MCP tools with budgets.
LLM
The model behind the agent.
IN THE FRAMEWORK
FoundryChatClient(project_endpoint, model, credential)
IN HAMBA
One client per agent. Model name from the environment. Terra to think, Sol to write, GPT-Image-2 to draw.
Read it bottom up. A model becomes an agent when you give it a role, a contract and tools. Agents become a team when code decides the order. The host makes the team callable.
henkboelman.com
Read slide text
LLMs
agents/brief_maker/main.py · agents/writer/main.py
from agent_framework.foundry import FoundryChatClient
from azure.identity import DefaultAzureCredential
credential = DefaultAzureCredential() # the agent's own identity
# brief-maker: one client, one model
client = FoundryChatClient(
project_endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"],
model=os.environ["BRIEF_MAKER_MODEL"], # gpt-5.6-terra
credential=credential,
)
agent = create_brief_maker_agent(client)
# writer: two clients, two prompts, no translation
english = FoundryChatClient(..., model=os.environ["ENGLISH_WRITER_MODEL"])
dutch = FoundryChatClient(..., model=os.environ["DUTCH_WRITER_MODEL"])
agent = create_writer_agent(english, dutch) # routes on language
# every Agent pins its output shape
default_options={"response_format": BriefPlan, "max_tokens": 6000}
1
2
3
1
The model is a client object.
FoundryChatClient talks to the project with the hosted agent's managed identity. The agent code never sees a key, an endpoint or a model name. All three arrive from outside.
2
The model is a variable.
BRIEF_MAKER_MODEL, ENGLISH_WRITER_MODEL, DUTCH_WRITER_MODEL. Swapping Terra for Sol on one role is an azd env set and a redeploy of one container. Nothing in the prompts changes.
3
Every call has a shape.
response_format is a Pydantic model on every agent. The LLM returns BriefPlan or ArticleDraft or ResearchQuestion, never a paragraph the next agent has to parse.
henkboelman.com
Read slide text
Models
Agents
What the job needs
Why this one
GPT-5.6 TERRA
brief-maker, researcher, editorial-reviewer, destination-linker, hotspot-writer
Holds a brief, twenty questions and a stack of sources. Judges, links, corrects.
The thinking jobs. Evidence in, structured verdicts out, one stateless pass per language.
GPT-5.6 SOL
writer (English), writer (Dutch), country-writer
Native prose in the magazine's voice. Two drafts, two prompts, no translation.
Neither language path ever sees the other draft. Separate voice files, shared brief.
GPT-IMAGE-2
image-creator
One hero, up to two heading-targeted inline images, private blobs.
Provisioned in the Foundry project, 30-minute SAS, imported into Hamba's media library.
NONE
travel-writer
Runs the workflow. Holds the job id, retries and stop conditions.
The orchestrator is code. No model owns workflow identity or draft metadata.
01
Config, not code.
Every model name is an azd environment variable. A swap is a settings change.
02
Structured output, always.
response_format=ArticleDraft. Contracts are Pydantic, not tagged prompt envelopes.
03
Two natives, no translator.
Dutch and English are written independently from one brief. Quality beats symmetry.
henkboelman.com
Read slide text
Agents
agents/researcher/agent.py
from agent_framework import Agent, MCPStreamableHTTPTool
from agent_framework import function_middleware
@function_middleware
async def enforce_web_research_limits(context, call_next):
# 10 searches, then 10 browses of URLs discovery returned
# 2 s between calls, 20 s timeout, one retry on 429
...
def create_researcher_agent(client) -> Agent:
web_iq = MCPStreamableHTTPTool(
url=os.environ["RESEARCHER_WEB_IQ_MCP_SERVER"],
header_provider=lambda _: {"x-apikey": os.environ["..."]},
allowed_tools=["web", "browse"], # not news, images, videos
request_timeout=20, approval_mode="never_require",
)
return Agent(
client=client, # the LLM, from main.py
instructions=RESEARCHER_INSTRUCTIONS, # prompt.md
name="researcher",
tools=[web_iq],
middleware=[enforce_web_research_limits],
default_options={"response_format": ResearchQuestion,
"max_tokens": 14000, "verbosity": "low"},
)
1
2
3
1
An agent is five arguments.
A client for the model, instructions from a file, tools, middleware, and default options with the output contract. That is the whole definition. Nine agents, one shape.
2
Tools are MCP servers, trimmed.
Web IQ exposes five tools. allowed_tools lets the researcher see two. The API key comes from a header provider reading the environment, never from the prompt or the code.
3
Middleware is where the rules live.
function_middleware wraps every tool call. It counts searches and browses, refuses a browse of any URL discovery did not return, paces calls two seconds apart and retries on 429. The model cannot negotiate with it.
henkboelman.com
Read slide text
Workflows
take_assignment
@step
create_brief
@step
research
@step
write_drafts
@step
finalize_drafts
@step
validate
@step
link_destinations
@step
create_images
@step
import_media
@step
save_story
@step
complete
@step
container restarts here
replays from the last checkpoint, skips research
A2A collaborator
Hamba MCP
code only
parallel EN ∥ NL, one step
01
Checkpointed by name and index.
Every @step result is persisted by Foundry hosting. A restart replays the function; finished steps return their saved result instead of running again.
02
Parallel work inside one step.
English and Dutch drafts run concurrently, but inside a single @step, so replay order is deterministic. The docstring in workflow.py says exactly this.
03
A workflow is an agent too.
workflow.as_agent(name="travel-writer") gives the whole function a name, a description and a session. The host cannot tell it from a chat agent.
henkboelman.com
Read slide text
Pydantic
define once, use three ways
from pydantic import BaseModel, Field, HttpUrl, ValidationError
class Source(BaseModel):
source_id: str = Field(pattern=r"^RQ\d+-S\d+$")
url: HttpUrl
title: str = Field(min_length=3, max_length=120)
Source.model_json_schema() # JSON Schema for the LLM
Source.model_validate_json(raw) # str/bytes in, Source out
# ValidationError: 1 validation error for Source
# title
# String should have at least 3 characters
# [type=string_too_short, input_value='AI']
source.model_dump_json() # Source in, bytes on the wire
# One class. Prompt schema, wire format, and the check on return.
1
2
3
1
Types are the schema.
A Python class with type hints. The rules live on the fields: patterns, lengths, ranges, URLs. There is no separate schema file to keep in sync.
2
The class becomes JSON Schema.
model_json_schema() gives the LLM a structured-output target. What you typed is exactly what the model is asked to produce.
3
Errors are data, not exceptions.
Bad input raises a ValidationError that names the field, the rule and the offending value. Readable by a human, and by the model on the retry.
Data validation for Python, driven by type hints. Shared by FastAPI, the OpenAI SDK and Agent Framework.
henkboelman.com
Read slide text
Web IQ
researcher → api.microsoft.ai/v3/mcp · one question
PHASE 1 · DISCOVER
web(query) ×≤10
Passages back, not pages: title, URL, snippet, timestamp, provenance. The model picks what deserves a closer read.
10 searches
10 passages each
2 s apart
PHASE 2 · READ
browse(url) ×≤10
Selected URLs fetched as Markdown, only from phase 1 results. Facts become evidence, never copy.
10 browses
20 s timeout
1 retry, 15 s
ResearchQuestion
answer · status · sources[] → ids validated in code, unresolved stops the run
1
Grounding built for agents.
Microsoft's replacement for the retired Bing Search APIs, announced at Build 2026. One MCP server, five tools: web, news, images, videos, browse. It returns evidence objects sized for a context window, not web pages.
2
Two phases, two budgets.
Search first, read second, and the second phase can only touch what the first returned. Both budgets sit in Agent Framework middleware, so a prompt cannot talk the researcher into a hundred calls.
3
Evidence, not copy.
Every answer carries sources whose ids are validated before drafting. Reports are treated as evidence, never as text to reuse. An unresolved question after one content retry stops the run instead of becoming plausible prose.
Limited access today: Web IQ is in preview for selected customers. Hamba authenticates with an API key that lives in the azd environment, never in source.
henkboelman.com
Read slide text
Contracts
agents/shared/contracts.py
ResearchQuestionId = Annotated[str, StringConstraints(
pattern=r"^RQ(?:[1-9]|1\d|20)$")] # RQ1 … RQ20, nothing else
class ResearchQuestion(BaseModel):
model_config = ConfigDict(extra="forbid") # unknown field = rejected
question_id: ResearchQuestionId
question: ResearchQuestionText # 20 to 1200 chars
status: ResearchStatus = ResearchStatus.PENDING
answer_markdown: str | None = None # max 12000 chars
sources: list[ResearchSource] = Field(max_length=10)
unresolved_reason: ShortString | None = None
@model_validator(mode="after")
def validate_research_state(self):
if self.status is ResearchStatus.UNRESOLVED:
if self.answer_markdown or not self.unresolved_reason:
raise ValueError("unresolved: a reason, no answer")
cited = set(EVIDENCE_MARKER_PATTERN.findall(self.answer_markdown))
if not cited:
raise ValueError("answer must cite at least one source ID")
if cited - {s.source_id for s in self.sources}:
raise ValueError("answer cites unknown source IDs")
1
2
3
1
One class, three jobs.
The same Pydantic model is the response_format the LLM must fill, the payload that crosses A2A, and the validator the orchestrator runs. Both ends import it from one file. Nothing is parsed twice.
2
The type system is part of the prompt.
Constrained strings, regex ids, extra="forbid". A question id is RQ1 to RQ20 or it does not exist. The model gets the schema; the rest is not up for interpretation.
3
Editorial rules live in validators.
An answer must cite a source it actually lists. Unresolved means a reason and nothing else. When validation fails, the error text goes back to the model with the retry, once.
Not tagged prompt envelopes. Shared classes, imported on both sides of every call, validated on the way back.
henkboelman.com
Read slide text
Agents talking to agents
henkboelman.com
Read slide text
Agent to Agent
submitted
working
input-required
auth-required
completed
failed
canceled
rejected
a task moves left to right. The two terracotta states pause and wait for the caller. The four dark ones are final.
CHAT
Part.text
Talk in turns.
Text parts on a message. The same taskId and contextId carry the conversation, and input-required lets the remote agent stop and ask you something.
STRUCTURED
Part.data
Send objects, not prose.
A JSON payload as a part. Each skill on the card declares inputModes and outputModes as MIME types, and the caller says which acceptedOutputModes it will take back.
ATTACHMENTS
Part.file
Files, either way.
Raw bytes inline or a url, with a mediaType and name. Images, PDFs, audio. In requests and inside artifacts coming back.
ARTIFACTS
Task.artifacts[]
Typed, named output.
Not a reply, a list of results, each made of parts. Delivered whole or in chunks with append and lastChunk while the task is still working.
STREAMING
message/stream
Watch it happen.
Server-sent events: TaskStatusUpdateEvent on every state change, TaskArtifactUpdateEvent as output lands. SubscribeToTask reattaches after a dropped connection.
PUSH
pushNotificationConfig
Get called back.
Register a webhook on the task. The remote agent calls your URL when the state changes. For work that outlives any connection, and for callers that cannot hold one.
Discovery is the Agent Card at /.well-known/agent-card.json: skills, modes, auth schemes, and whether it streams or pushes. JSON-RPC, gRPC or plain HTTP underneath.
henkboelman.com
Read slide text
The payload
POST / · message/send · A2A-Version: 1.0
{
"jsonrpc": "2.0",
"id": "req-1",
"method": "message/send", // or message/stream for SSE
"params": {
"message": {
"role": "user", // replies come back as agent
"messageId": "m-01",
"contextId": "ctx-rome-42", // groups related tasks
"taskId": null, // null = new task; set it to continue
"parts": [
{ "text": "Check these claims." }, // chat
{ "data": { "claims": ["Trevi is 26 m high"] } }, // structured
{ "file": { "url": "https://…/draft.md", // attachment
"mediaType": "text/markdown" } } // or raw base64
]
},
"configuration": {
"acceptedOutputModes": ["application/json"], // what I accept back
"returnImmediately": false, // block until done or paused
"historyLength": 0 // do not echo the conversation
}
}
}
200 OK · result: Task
{
"jsonrpc": "2.0",
"id": "req-1",
"result": {
"id": "task-7f3a", // minted by the remote agent
"contextId": "ctx-rome-42", // same conversation
"status": {
"state": "TASK_STATE_COMPLETED", // or INPUT_REQUIRED …
"timestamp": "2026-09-03T09:12:44Z"
},
"artifacts": [ // the output
{
"artifactId": "a-1",
"name": "verdicts",
"parts": [ // same three part kinds
{ "data": { "verdicts": [
{ "claim": "Trevi is 26 m high",
"verdict": "true",
"source": "https://…" } ] } }
]
}
],
"history": [] // historyLength was 0
}
}
henkboelman.com
Read slide text
Parts
TEXT
{
"text":
"Check these claims.",
"mediaType": "text/markdown"
}
For the model.
Plain text or markdown. The instruction, the question, the answer a person will read. Any mediaType you like.
DATA
{
"data": {
"question_id": "RQ3",
"claims": [ "Trevi 26 m" ]
},
"metadata": { "schema": "RQ" }
}
For the code.
Any JSON value: object, array, string, number. A contract that travels as a contract, with its own media type.
FILE · URL
{
"url":
"https://…/brief.md",
"mediaType": "text/markdown",
"filename": "brief.md"
}
A pointer.
The receiver fetches it. Big things, shared storage, a short-lived SAS link. Nothing heavy on the wire.
FILE · RAW
{
"raw":
"iVBORw0KGgoAAAANS…",
"mediaType": "image/png",
"filename": "hero.png"
}
The bytes.
Base64 inline. Small things, or when the receiver cannot reach your storage. mediaType is required.
Mix them in one message. Every part may carry mediaType, filename and metadata. The caller says what it will take back with acceptedOutputModes. Artifacts use the same four kinds, so what comes back looks like what went in.
Hamba sends one text part: the contract, serialized. Both ends already import the same class.
henkboelman.com
A2A in Hamba
agents/shared/a2a.py
async def discover_agent(name, project_endpoint, http):
url = f"{project_endpoint}/agents/{name}/endpoint/protocols/a2a"
card = await A2ACardResolver(http, url).get_agent_card(...)
return A2AAgent(agent_card=card, http_client=http, timeout=300)
class AzureBearerAuth(httpx.Auth): # managed identity, no keys
token = credential.get_token("https://ai.azure.com/.default")
class StructuredA2AClient:
async def invoke(self, request, response_model, max_attempts=2):
task = request.model_dump_json() # Pydantic in, as text
with tracer.start_as_current_span(f"invoke_agent {self.name}"):
_record_payload(call_id, "request", task) # exact, hashed
response = await self.agent.run(task, session=session)
state = _task_state(session.service_session_id)
if state != TaskState.TASK_STATE_COMPLETED: # else: fail
errors.append(f"task_state={state}")
elif response.user_input_requests: # no questions
errors.append("interactive input requested")
else:
return response_model.model_validate( # Pydantic out
_decode_json(response.text))
task += "Return the corrected JSON. Previous error: " + errors[-1]
1
2
3
1
Find it by card, call it by identity.
Foundry publishes every hosted agent's card under the project. travel-writer resolves it by name and calls with a managed-identity bearer token. Three hundred seconds per task, because research is slow.
2
Structured, over the text part.
A Pydantic request goes out as JSON in a text part. The remote agent runs with response_format set to the reply contract. The reply text is validated back into Pydantic. Typed both ways, on the one part every client supports.
3
Only COMPLETED counts.
Any other task state is an error with the state name on the span. A collaborator that asks a question is treated as a failure. Bad JSON gets one retry with the validation error appended. Then the workflow decides.
Switched on card discovery, structured payloads, artifacts, task state. Left off streaming, push, input-required, file parts. On purpose.
henkboelman.com
Read slide text
Running it
henkboelman.com
Read slide text
Where to host
MICROSOFT FOUNDRY · HOSTED AGENTS
hamba-agents-ai-team
travel-writer · workflow host · country-writer
A2A · agent card · managed identity bearer token
brief-maker
researcher
writer
editorial-reviewer
destination-linker
hotspot-writer
image-creator
gpt-image-2
nine containers, python 3.13, 0.5 cpu / 1 GiB each · azd deploy <agent> · collaborators first, workflow host last
HAMBA WEBSITE
hamba.nl/mcp
.NET + SQL + MCP server. Assignments in, drafts, media and completion out. Direct call_tool, never LLM-selected.
WEB IQ
api.microsoft.ai/v3/mcp
Research. Ten searches, ten browses per question, serialized and paced by middleware.
PRIVATE BLOB STORAGE
generated-images
JPEG blobs, 30-minute user-delegation SAS, then imported into Hamba's media library.
MCP
MCP
SAS
azd provision once for the project, image model and storage. Then azd deploy per agent. Secrets live in the azd environment, never in source.
henkboelman.com
Read slide text
Foundry Hosted Agents
azure.yaml · brief-maker
brief-maker:
host: azure.ai.agent
kind: hosted
language: python
codeConfiguration:
entryPoint: agents/brief_maker/main.py
runtime: python_3_13
container:
resources: { cpu: "0.5", memory: 1Gi }
env:
BRIEF_MAKER_MODEL: ${BRIEF_MAKER_MODEL}
protocols:
- protocol: responses
version: 2.0.0
agentEndpoint:
protocols: [responses, a2a]
agentCard:
description: Extracts factual coverage while preserving
assignment language and content intent.
version: "1.0"
skills:
- id: story-briefing
name: Content briefing
1
2
3
1
Your code, their infrastructure.
A hosted agent is your own container on Foundry Agent Service. Any framework. You pick CPU and memory, the platform runs it in a per-session VM sandbox that scales to zero and persists state between turns.
2
It speaks two protocols.
Responses for clients and the portal. A2A so other agents can find it by its card and hand it a task. Each version gets its own endpoint under the project.
3
Identity comes with the deploy.
Every agent gets a dedicated Entra ID when it ships. That is how nine agents call each other, the models and the storage account with no key in any container.
You own the code inside the sandbox. The platform owns the endpoint, identity, scaling and session state around it.
henkboelman.com
Read slide text
Toolbox
toolbox.yaml · main.py
# toolbox.yaml azd ai toolbox create hamba-tools --from-file .
description: Every tool a Hamba agent may touch, one endpoint
tools:
- type: mcp # the website
server_label: hamba
project_connection_id: hamba-mcp-conn
require_approval: "never"
- type: mcp # Web IQ, key in the connection
server_label: web-iq
project_connection_id: webiq-conn
- type: image_generation # built-in Foundry tool
# main.py
from agent_framework.foundry import FoundryChatClient, FoundryToolbox
toolbox = FoundryToolbox(credential) # TOOLBOX_ENDPOINT from env
agent = Agent(client=client, instructions=..., tools=toolbox)
1
2
3
1
One endpoint, every tool.
A toolbox is a versioned bundle of tools, published as a single MCP endpoint in the Foundry project. Built-in tools, MCP servers, OpenAPI, A2A. Agents connect once.
2
Keys leave the agent.
Today the Web IQ key is an environment variable in the researcher's container. In a toolbox it lives in a project connection; the agent brings its own identity.
3
Two lines in the agent.
FoundryToolbox(credential) resolves the endpoint from the environment and forwards the platform call-id. tools=toolbox. Same Agent class, no per-tool wiring.
Generally available in Foundry. Not in hamba yet: it would replace the per-agent wiring on the Agents slide.
henkboelman.com
Read slide text
How to monitor
Hamba Trace Visualizer · operation 9c41…e2 · exact
hamba-editorial-workflow
take_assignment · MCP
create_brief · A2A
research ×6 · A2A
web_iq.search / browse
write_drafts EN ∥ NL
finalize_drafts EN ∥ NL
validate · code
destination-linker · A2A
image-creator · A2A
save + complete · MCP
1
Spans come from the framework.
Agent Framework emits OpenTelemetry. a2a.py adds a span per collaborator call with the task state, and Foundry Observability shows the same run from the other side.
2
Payloads are recorded exactly.
Every A2A request and raw response is chunked into telemetry, length and SHA-256 checked on replay. Runs without it are labelled inferred, never dressed up.
3
One id replays the whole run.
Paste an Application Insights operation id into the visualizer on Container Apps. Handoffs, the brief, every Web IQ search, every source, both drafts.
henkboelman.com
Read slide text
Demo: Hamba
hamba.nl · assignments
SCREENSHOT
The assignment
IN
One line from an editor, in Dutch or English. A writer personality if you want one.
hamba-trace · operation 9c41…e2
SCREENSHOT
The run
DURING
Nine agents, one operation id. Every payload that crossed the wire, as it happened.
hamba.nl · the article
SCREENSHOT
The article
OUT
Two languages, linked destinations, hero and inline images. Waiting for one human click.
Start the run before the A2A section. By now it is done.
henkboelman.com
Read slide text
Lessons learned
01
Control flow in code, not in a prompt.
The travel-writer is a @workflow, not an LLM coordinator. Every transition is Python you can read, checkpoint and replay.
02
Contracts before prompts.
One Pydantic class per boundary. Half the rules moved out of prompt.md into types the model cannot argue with.
03
One job, one agent, one model.
Terra where it has to think, Sol where it has to write, Image-2 for pictures. The model is a variable, not an identity.
04
Switch on the protocol you need.
Card discovery, structured payloads, task state. Streaming and push stayed off until something asked for them.
05
Record what crossed the wire.
Exact payloads with a hash, replayable by one operation id. You cannot fix a run you cannot see.
06
Validate in code, retry once.
The reviewer gets one correction pass with the exact error. Then the step decides, not the model.
None of it is about the model. All of it is about the seams between them.
henkboelman.com
Read slide text
Thank you.
Your new teammates. Architecting a multi-agent workforce.
hamba.nl the magazine
henkboelman.com slides and the write-up
henkboelman.com
in/henkboelman
hnky
youtube.com/henkboelman
@hboelman
Slide 1 of 51
Use Previous and Next, choose a slide, or use the left and right arrow keys while the viewer is focused.