Huge thanks to everyone that helped us pull off the 32-hour livestream last week. Good news is the live stream recordings are up now. Check them out in case you want to relive some moments or catch a talk you may have missed.
Day one and day two |
In case you missed it GemsHuman validation for a news bias model Built around the open-source BiasExpert model, the study asks readers to assess bias in short news passages, helping measure where AI-generated classifications align with human judgments. Technical post on compositional agent harnesses Training an RLM harness on short tasks generalized to held-out tasks 8-32x longer, offering agent builders a concrete approach to context offloading, programmatic sub-agents, and cross-domain generalization. Agent engineering playbook toolkit Built around 23 playbooks and modular skills, pstack gives coding agents structured workflows for debugging, verification, architecture, parallel work, reviews, and long-running engineering tasks across multiple agent environments. Paper on cheaper agent self-improvement Using pairwise LLM judging and Bradley-Terry ranking, SIFT steers tree search toward promising code modifications while reserving expensive benchmark evaluations for stronger candidates, reducing compute and evaluation costs. |
LOCAL ORGANIZERS What’s happening across the chapters |
AGNTCON + MCPCON NORTH AMERICA Engineering the agentic stackJoin developers, maintainers, engineering leaders and open-source contributors in San Jose on October 22–23 for two days focused on the systems behind agentic AI. Expect technical talks, implementation lessons and conversations around MCP, agent infrastructure, orchestration, security, observability, evaluation and the open standards shaping how agents are built and operated. It’s also a chance to meet the people behind the projects, compare approaches with teams working on similar problems, and come away with ideas you can apply to your own stack. Use code COMMUNITY25 for 25% off registration. Register for AGNTCon + MCPCon North America → |
LUNCH AND LEARN Verification and recovery What happens after a verifier says no?Dolly Sah and Tanmay Sah’s Session 25 looked at what happens when agents can identify unsafe actions but struggle to recover from them. The write-up covers the verifier tax, unsafe success, and EvoUndo’s approach to making agent self-modification reversible. Read the write-up here. This week’s Coding Agents Lunch & Learn is Friday, September 25 at 9 AM PT / 6 PM CEST. Session 26 drops the usual featured talk for an open discussion on coding agents - what people are building, which tools and workflows are working, where agents still fall short, and the challenges around context, verification, reliability and autonomy. Join the next session here. |
READING GROUP Can agents prove they followed policy? The next AAIF Reading Group looks at Measuring Agent Policy Compliance, a paper on evaluating whether AI agents follow their intended policies. The group will work through the paper’s methodology, assumptions and limitations, with plenty of room to question the approach and discuss what it means for agent evaluation. Thursday, October 1 · 9 AM PT / 6 PM CEST Join the reading group here. |
AAIF COMMUNITY Walking Tokyo Talking Agent ProtocolsBrowser agents burn tokens taking screenshots, searching accessibility trees, and clicking through pages they barely understand. WebMCP offers a cleaner path by exposing page-specific tools inside the user’s authenticated session, alongside agent commerce and A2A. - In China, super-app ecosystems reduce integration overhead, but expansion beyond them increases the value of MCP and A2A.
- WebMCP could make browser automation faster and more reliable by replacing visual interaction with contextual tools tied to the page.
- Agentic commerce still needs identity, authorization, and payment mechanisms, with KYC extending to agents acting for customers.
The shift is toward agents operating through explicit tools, permissions, and protocols rather than brittle UI automation. Video · Spotify · Apple |
Voice Agent - Virtual EventA semantic cache cut one voice-agent response from 1.1 seconds to 68 milliseconds in a live demo. Across the event, the recurring challenge was making real-time agents fast, controllable, and robust outside clean test environments. - Breaking monolithic prompts into scoped steps more than halved token use while making required checks enforceable.
- Production latency depends on the whole stack: model choice, network distance, caching, turn detection, and prefetching.
- Real-world reliability needs noisy-audio testing, P50/P95/P99 monitoring, human escalation, and simulations that target failures such as interruptions, transcription errors, and privacy leaks.
Strong voice systems come from controlling context, infrastructure, audio, and evaluation together. Watch the event |
Introducing the First Open Source Vending Machine BenchmarkA green “completed” result can still hide an agent that loses money. This benchmark turns a simulated vending business into a controlled environment for testing how tool-using agents handle pricing, inventory, purchasing, refunds, and changing demand. - Six machines across three locations expose agents to shared cash, limited working time, supplier failures, and demand shifts.
- The simulator uses 22 MCP tools and reproducible rules for weather, seasonality, pricing, marketing, and customer preferences.
- Early runs showed large differences in business outcomes between models, while higher API spend did not consistently produce better results.
The useful signal is whether an agent can make sound operating decisions repeatedly, not simply finish the task. Read the blog |
402 payment required: what enterprise MCP servers owe the agents that pay themA paid MCP tool creates more than a billing problem: the agent can choose the work, but it should not control the budget that approves it. x402 v2 and MCP 2026-07-28 make pricing and routing metadata more explicit, which shifts attention to where enforcement sits. - Servers should recompute prices from the current request, rather than trust values supplied by the client.
- Replay protection, local payee checks, and independent budget controls become part of the execution path.
- HTTP and MCP payment bindings expose different tradeoffs for gateways, clients, and shared infrastructure.
The core design question is how to separate agent intent from the authority to spend. Read the blog |
|
Come and connect IN-PERSON EVENTSFind your city here, or start a chapter if there isn't one yet. |
Join from anywhere VIRTUAL EVENTS |
MEME OF THE WEEK  |
Sent by Demetrios Brinkmann on behalf of Agentic AI Foundation (AAIF). The Linux Foundation, 2810 N Church St., PMB 57274, Wilmington, Delaware 19802-4447, United States |
|
|