![]()
Di previous lesson scale all di agents up go cloud. Dis one dey bring dem down go single machine. By di time you finish, you go get one working engineering assistant wey go reason, call tools, read your files, and search your documentation — without make one single cloud inference call.
Why you go want am? Three reasons wey dey come up anyhow for real engineering work:
Di problem be say you dey exchange one frontier cloud model for Small Language Model (SLM) wey dey run on your CPU, GPU, or NPU. Dis lesson na about to build agents wey go good inside dat kind limit instead of to pretend say di limit no dey.
Dis lesson go cover:
After you finish dis lesson, you go sabi how to:
Dis lesson assume say you don finish di earlier lessons and you sabi:
You go also need:
requirements.txt, plus foundry-local-sdk, openai, and chromadb.One frontier cloud model get hundreds of billions parameters and e get one big data centre behind am. SLM get small billion parameters and e for fit your laptop RAM. Dis difference dey set correct expectation.
SLMs good for:
SLMs no too strong for:
Di best way for local agents na: make SLM dey control, tools dey carry heavy work. Di model no need to know your codebase — e need sabi when to call read_file and search_docs. Na wetin SLM good for.
flowchart LR
U[Developer] --> A[Local SLM Agent]
A -->|dey choose which tool| T1[read_file]
A -->|dey choose which tool| T2[search_docs RAG]
A -->|dey choose which tool| T3[analyze_code]
T1 --> A
T2 --> A
T3 --> A
A --> R[Answer, fully on-device]
Microsoft Foundry Local na lightweight runtime wey dey download, manage, and serve models fully on your machine. Wetin important for us na say e get OpenAI-compatible HTTP endpoint — dat one mean OpenAI SDK and Microsoft Agent Framework’s OpenAI client fit work for am by just changing base_url. Wetin you don learn on how to build agents, e fit work same way; na only di endpoint go move from cloud go localhost.
Foundry Local go select di best build of model for your hardware automatically — CPU build, CUDA/GPU build, or NPU build — no need to hand-optimize for each machine.
Install Foundry Local (check documentation for your OS), then confirm say e dey work:
# Install (for example; follow di docs for your platform)
winget install Microsoft.FoundryLocal # Windows
# brew install microsoft/foundrylocal/foundrylocal # macOS
# Download an run one Qwen model, den start di local service
foundry model run qwen2.5-7b-instruct
foundry service status
Once di service dey run you get local OpenAI-compatible endpoint (normally http://localhost:PORT/v1). Di notebook dey use foundry-local-sdk to find di endpoint automatically, so you no need hard-code di port.
Agent na agent only if e fit call tools. Plenti SLMs fit chat but dem no fit produce reliable, correct tool calls. Qwen models train to do function calling well and dem dey produce correct tool call structures steady — na wetin turn local chat model to local agent.
Di flow na di normal tool-calling loop wey you sabi, but e dey run for inside device:
sequenceDiagram
participant U as User
participant A as Qwen Agent (local)
participant T as Local Tool
U->>A: "Wetín auth.py dey do?"
A->>A: Decide: call read_file
A->>T: read_file("auth.py")
T-->>A: file contents
A->>A: Reason over contents
A-->>U: Explanation
Documentation search na where local agents show their work. Instead of hope say di SLM don memorize your framework docs, you go embed docs for local vector database and make agent retrieve correct parts anytime e need am.
We dey use Chroma, one embedded vector store wey dey run inside process with no server. Di pipeline na local: local embedding model → local vectors → local retrieval → local SLM.
flowchart TB
D[Your docs / code] --> E[Local embedding model]
E --> V[(Chroma vector DB - on disk)]
Q[Agent query] --> QE[Embed query locally]
QE --> V
V -->|top-k chunks| A[Qwen agent]
A --> Ans[Grounded answer]
Dis na di same Agentic RAG pattern from Lesson 5 — only difference na say every part dey run on your machine.
MCP no be cloud service, na transport. MCP server fit run as local process on stdio, expose tools to your agent with standard protocol. E make you fit reuse di plenti MCP servers — filesystem access, git operations, database queries — fully offline.
Security no be like cloud, but e no mean say e no get security: local MCP server dey run with your user permission, so limit wetin e fit touch (for example, only project directory, no be your whole home folder) and always check outputs before make use.
Local first no mean na only local. Mature systems go select path based on sensitivity and difficulty:
| Situation | Where e go run |
|---|---|
| Sensitive code/data or offline | Local SLM |
| Simple, bounded task | Local SLM (cheap, fast) |
| Hard multi-hop reasoning on non-sensitive data | Cloud model |
| Everything during outage | Local SLM (graceful degradation) |
Dis dey similar to model routing idea from Lesson 16 — only difference na one of di “models” na your own machine. Good design go fallback to local when cloud no dey, so agent no go fail but e go just reduce quality small.
flowchart LR
Q[Request] --> S{Sensitive or offline?}
S -->|yes| L[Local SLM]
S -->|no| C{Need deep tink?}
C -->|no| L
C -->|yes| Cloud[Cloud model]
L --> Out[Response]
Cloud --> Out
Open code_samples/17-local-agent-foundry-local.ipynb and follow am. You go build local engineering assistant wey go run fully on your workstation and e fit:
No cloud inference dey anywhere.
Agent connect to Foundry Local via OpenAI-compatible endpoint, so di agent code close to cloud lesson code — na client part change:
from foundry_local import FoundryLocalManager
from openai import OpenAI
# Foundry Local sabi/find di model and e give us local endpoint.
manager = FoundryLocalManager(\"qwen2.5-7b-instruct\")
client = OpenAI(base_url=manager.endpoint, api_key=manager.api_key) # api_key na local placeholder.
Tools na normal Python functions wey scoped to project directory:
def read_file(path: str) -> str:
\"\"\"Read a file, but only inside the sandboxed project directory.\"\"\"
full = (PROJECT_ROOT / path).resolve()
if PROJECT_ROOT not in full.parents and full != PROJECT_ROOT:
return \"Access denied: path is outside the project directory.\"
return full.read_text(encoding=\"utf-8\")
Remember sandbox check — even for local, tool wey read random path fit cause wahala. Di notebook keep every tool scoped to one project root.
Test yourself before you go to assignment.
1. Give me two real reasons to run agent locally instead of for cloud.
2. How dem recommend to divide work between SLM and tools for local agent, and why?
3. Wetin make you fit reuse cloud agent code with Foundry Local?
4. Why you choose Qwen function-calling model and no any SLM?
5. For local RAG pipeline, which parts run for machine?
6. Local MCP server dey run for your machine. E mean say e automatic safe? Wetin you still go do to stay safe?
7. Talk one correct hybrid routing rule wey include local model?
8. Na how many minimum RAM wey dey realistic to run local agent for dis lesson, and wetin more RAM fit give you?
Extend local engineering assistant to be local documentation reviewer for small project wey you choose (fit use one of dis repo lesson folders).
Your submission suppose:
Add find_todos tool wey go scan project for TODO/FIXME comments and return dem with file and line number — also keep sandbox check same as read_file.
Den write one short paragraph about wetin you go move go cloud and wetin you go keep local for dis reviewer, plus why. Dem go check if di local parts tie together well and if your hybrid reasoning correct — no be about model quality.
For dis lesson you build one agent wey dey run fully for your own machine:
Dis one complete di deployment journey: Lesson 16 scale agents up go Microsoft Foundry, and dis lesson scale am down for one single workstation. Di next lesson go show how to keep deployed agents safe.
Disclaimer: Dis document don translate wit AI translation service Co-op Translator. Even tho we dey try make am correct, abeg make you know say automated translation fit get errors or mistakes. Di original document for dia own language na im be di correct source. For important info, make person wey sabi human translation do am. We no go responsible for any misunderstanding or wrong understanding wey fit happen because of dis translation.