Depending on when you stop for your coffee break to read this, we might already be live from Amsterdam. We’re broadcasting AGNTCon + MCPCon Europe for 36 hours straight - streaming the talks, wandering the expo floor, grabbing speakers and builders for chats, keeping things going after the conference closes, and seeing what happens through the night. Come hang out with us over the next couple of days! |
In case you missed it GemsOpen control plane for agent operations Built around Kubernetes and OpenTelemetry, it brings deployment, lifecycle management, policy enforcement, observability, and external-agent monitoring into one place for teams operating agents at scale. Agent harness red-team benchmark Testing five models across 1,000 attacks, the comparison isolates how Claude Agent SDK and deepagents affect attack success, prompt scaffolding, tool behavior, and self-reported red-team results. Agent memory poisoning defense study Using LongMemEval, the paper tests content screening and provenance-weighted retrieval against poisoned memories, finding both can fail and proposing bounded retrieval constraints as a more robust defense. Agent environment and eval task guide Turning traces, code, and human input into reusable specs, the workflow separates task design from implementation to help teams build realistic agent benchmarks and continuously refine them. |
LOCAL ORGANIZERS What’s happening across the chaptersThree more AAIF chapters held their first meetups last week, with Seoul, Toronto and Colombo all bringing new local communities together. Toronto covered evals for Kubernetes agents, enterprise control planes and model routing, including one Jira example that cut input tokens from 13,514 to 1,025. Colombo ranged from MCP auth and graph-based grounding to WebMCP voice agents and a tool-design approach that cut one MCP server by 40%. Chennai also kicked off its chapter with a strong production focus, including hybrid retrieval, permission-aware search, OpenSearch observability and tracing the full prompt lifecycle with Jaeger and Langfuse. There’s more taking shape too: first events are on the way in Shenzhen and Bengaluru, while Agentic Tokyo #2 brought the local community together with AAIF visitors for another round of agents, open source and shared lessons. |
LUNCH AND LEARN Verification and recoveryLast week, Eric Bigelow from Goodfire joined us to look at uncertainty in LLM reasoning, including the points where a reasoning chain can suddenly change course and how to find them with far fewer samples. Read the session notes This week, Tanmay Sah and Dolly Sah join us on September 18 at 09:00 PDT / 18:00 CET to talk about verification and recovery in coding agents, from the tradeoffs of checking an agent’s work to recovering after a harmful or incorrect change. Join us |
NEW FROM AAIF The first official MCP certification is liveAAIF and Linux Foundation Education have released the Model Context Protocol Associate (MCPA), a vendor-neutral certification based on the 2026-07-28 MCP spec. Candidates answer questions on client-server interactions and the tool invocation lifecycle, with 24% of the exam focused on security and governance. See exam details |
AAIF COMMUNITY Why cost per million tokens is a useless KPIAgentic AI can turn predictable usage into runaway spend: each retry may trigger more model calls, RAG lookups, vector retrieval, memory, and egress. That makes cost per million tokens a weak KPI, because identical token volumes can represent very different workloads and value. - Cost models are shifting toward adoption rate and cost per user, with separate baselines for engineering, product, and internal use cases.
- Real-time gateways can cap retries, route models, apply fallbacks, and enforce cost and security policies before loops spiral.
- Value needs separate metrics, from DORA-style delivery measures to hours saved or support workload reduced.
The better unit is cost tied to a specific persona, workload, and outcome. Video · Spotify · Apple |
Same door. Very different walk in.An AI coding agent reported every unit test passed - because the tests were mocked so thoroughly that nothing meaningful ran. That failure frames a broader problem: agents can retrieve facts without preserving the context and relationships needed to use them safely. - Nexus stores semantic and episodic memory together, combining graph and vector search with multimodal data.
- Each memory commit carries who, what, when, why, and how, while lineage and query-time permissions remain separate trust requirements.
- Open gaps include recency-aware updates, permission enforcement, GraphRAG, hybrid search, and propagating corrected context across agents.
The architecture shows why persistent agent memory needs more than embeddings: relationships, provenance, governance, and update semantics all matter. Read the blog |
Dex still thinks long-running coding agents stinkCoding agents can ace bounded tasks and still wreck a codebase over time. In SlopCodeBench, every tested model accumulated defects across successive requirements, while a lightly supervised internal experiment ended with 30,000–40,000 lines of code being archived. - Across six challenges and 30 checkpoints, Fable and Sol each managed 10 strict passes, or 33.3%.
- The failure mode is compounding “slop”: agents inherit weak patterns from earlier code and turn them into precedent.
- Human review remains the bottleneck because architecture and “taste” failures are harder to verify than syntax, tests, or task completion.
The key challenge is building stronger verifiers that can catch quality decay before it becomes architecture. Read the blog |
From supervised shells to purpose-built MCP toolsA client asked an MCP server to run shell_exec and got an unknown-tool error - exactly the boundary a general-purpose shell cannot provide. The test used a Tesla T4 to show how a purpose-built GPU server can narrow agent access without pretending the tool schema solves everything. - Four typed GPU operations exposed inventory, metrics, summaries, and process data; invalid selectors and unregistered tools were rejected.
- Under CUDA load, utilization reached 100%, while NVML and PyTorch reported different memory totals because they measure different things.
- Least privilege still depends on process identity, device visibility, environment, transport, and logging.
The practical lesson is to narrow the MCP contract, the runtime process, and the evidence trail together. Read the blog |
|
Come and connect IN-PERSON EVENTSFind your city here, or start a chapter if there isn't one yet. |
Join from anywhere VIRTUAL EVENTS |
MEME OF THE WEEK  |
Sent by Demetrios Brinkmann on behalf of Agentic AI Foundation (AAIF). The Linux Foundation, 2810 N Church St., PMB 57274, Wilmington, Delaware 19802-4447, United States |
|
|