HIT REPLY “What do you check before switching models that isn’t in anyone’s eval suite?” Reply to this email - we read every one. |
In case you missed it GemsClaude Code token-routing guide Using Spotify Portal’s AiKA modes and a Claude Code plugin, the setup routes bulk reads and predictable code generation to cheaper models, cutting Claude token use by around 90% in benchmarks. Claude Code Function Hooks preview Still pre-release, the proposal adds TypeScript-based hooks for intercepting behavior, modifying UI, enforcing admin controls, and observing events, giving plugin authors a concrete look at the proposed extension model. Devbox platform for coding agents Built around full VMs rather than basic sandboxes, Namespace gives coding agents codebases, dependencies, databases, network controls, Mac/Linux environments, and interactive debugging for more complete development workflows. Security response plan for AI attacks Focused on cheap open-weight models automating vulnerability discovery and exploitation, this lays out practical defenses around patch deployment, scoped agent access, fuzzing, supply-chain inventory, containment, and recovery. |
FREE VIRTUAL EVENT Building the next voice agents What does it take to move voice AI from a slick demo to something fast, reliable, and ready for production? Join us September 16 for 90 minutes on voice-agent architecture, infrastructure, and open standards, including UNMUTE, a new MIT-licensed standard for voice agents, plus a live look at how semantic caching, CDNs, and edge infrastructure can cut latency in voice pipelines. September 16 · 08:30 PDT / 17:30 CEST JOIN LIVE |
READING GROUP Prompt Injection as Role ConfusionThis month’s reading group looked at research showing that models can infer who is speaking from writing style rather than reliably following role tags. With forged reasoning written in a model’s own style, attack success rose from under 4% to over 80% on several of the six models tested. Lucas Pavanelli walked through the ICML 2026 paper, before Sparsh Jain took the discussion into reasoning-trace extraction and agent red-teaming. Read the session notes and checklist |
LOCAL ORGANIZERS What’s happening across the chaptersIt’s great to see AAIF chapters taking shape around the world, with local organizers bringing people together and trying out their own formats. Melbourne’s first AAIF meetup brought 58 engineers and builders together for four talks. Pittsburgh’s hands-on MCP session drew around 40 people and so many questions that the agentgateway and A2A workshops were saved for another event, while Shanghai opened its Launch Series and Atlanta tried a speaker-free Lean Coffee format where the room chose the agenda.
Over in Silicon Valley, Vijay Bohre and Rahul Parundekar brought together three technical sessions on agent memory, covering everything from what should be remembered or forgotten to memory scoping, retrieval and keeping useful context under control. Watch it here. The community map is growing too, with five new organizers joining across Madrid, Santiago, Chennai and Ulaanbaatar. There’s plenty more happening this month - check out the upcoming events near the bottom of the newsletter. |
LUNCH AND LEARN Coding agents, security and uncertainty Last week, Ads Dawson joined us to talk about the security scaffolding around coding agents, from least privilege and sandboxing to prompt injection and stopping conditions. Read the session notes This week, Eric Bigelow from Goodfire joins us on September 11 at 09:00 PDT / 18:00 CET to look at uncertainty in LLM reasoning, and how to get useful signals from fewer sampled reasoning chains. Check out his paper on Model Reasoning, and bring your questions! Join us |
AAIF COMMUNITY The Five-Layer Cake Approach to Scaling AI Without Wasting MoneyThirty minutes of GPU warm-up across thousands of accelerators can waste expensive capacity. AI efficiency depends on choices across hardware, capacity, inference, models, and routing, where gains and mistakes compound. - Match hardware to workload: coding favors heavy prefill and context processing; multimodal pipelines may need different resources for vision and text.
- Treat caching, quantization, fleet partitioning, and engine startup as tradeoffs between throughput, flexibility, and utilization.
- Benchmark changes with repeatable workloads: model swaps and lower precision can improve speed, but they can also change output quality and behavior.
Stable baselines make it easier to turn efficiency gains into lower latency, more capacity, or larger reasoning budgets. Video · Spotify · Apple |
LiteLLM is the known option. agentgateway is the open one.AI gateways can sit in front of model keys, MCP servers, and A2A traffic, making licensing and failure modes architectural concerns. Comparing LiteLLM and agentgateway shows how differently they behave once identity, routing, and operations matter. - LiteLLM has broader provider coverage, but SSO, OIDC/JWT, secret managers, and governance controls require Enterprise.
- agentgateway keeps JWT, CEL, mTLS, MCP OAuth, and external authorization in its Apache 2.0 tree, with a smaller runtime footprint.
- Their attack surfaces differ: LiteLLM production deployments depend on PostgreSQL and Redis, while agentgateway can run as one binary.
The better default depends on whether provider breadth or an open, general-purpose control plane is the main constraint. Read the blog |
MCP 2026 changes the rulesMCP’s 2026-07-28 revision removes protocol-level sessions, changing how clients scale across load balancers and replicated servers. Requests now carry the metadata they need, while discovery and optional capabilities move out of the old initialization flow. - Stateless requests let servers process calls independently and scale horizontally without shared session state.
- Versioned extensions keep the core smaller, so features such as Tasks or MCP Apps can evolve separately and be adopted only where needed.
server/discover lets clients inspect supported versions and capabilities without creating a session or sending notifications/initialized.
The main shift is toward a more scalable, modular protocol with less connection state and clearer capability negotiation. Read the blog |
Authorization in MCP 2026-07-28: Clarifying trust relationships for agentic systemsAn MCP client may cross several authorization servers, backend services, and organizational boundaries in a single workflow. MCP 2026-07-28 tightens how those trust relationships are identified, validated, and maintained using established OAuth and OIDC mechanisms. - Clients now validate the
iss parameter when present, reducing the risk of accepting authorization responses from the wrong authority. - Dynamic registration requires an appropriate OIDC
application_type, while credentials remain bound to the authorization server that issued them. - Step-up authorization preserves existing scopes, and refresh-token guidance better supports long-running agent workflows.
The key change is clearer, more verifiable authorization across distributed MCP systems without requiring a new IAM model. Read the blog |
|
Come and connect IN-PERSON EVENTSFind your city here, or start a chapter if there isn't one yet. |
Join from anywhere VIRTUAL EVENTS |
MEME OF THE WEEK  |
Sent by Demetrios Brinkmann on behalf of Agentic AI Foundation (AAIF). The Linux Foundation, 2810 N Church St., PMB 57274, Wilmington, Delaware 19802-4447, United States |
|
|