It’s nearly Friday, and you know what that means? Yeah, the weekend, but more importantly, another Lunch & Learn session! This week, Satyam Soni is bringing agents.md into the mix, looking at what happens when coding agents can leave useful knowledge behind for the next task. Come and join us. |
In case you missed it GemsDoorDash’s Data Agent Architecture and Evals Built around retrieval, ranking, a governed semantic data layer, and 1,900-plus evals, Vera offers a concrete production example of improving agent accuracy while managing latency, tool use, and data quality. Lessons From Building Codex at OpenAI Drawing on work across Codex Web, CLI, Slack, code review, and GPU capacity, Kiriti Badam shares practical lessons on agent interfaces, reliability, context engineering, evals, and product iteration. Agent Plugins Standard for Portable Extensions Combining Agent Skills, MCP servers, manifests, and client-specific extensions in one package format, the specification gives authors and client implementers a shared interoperability layer across agent tools. Pi’s MCP and Codemode Architecture Pairing MCP with a JavaScript sandbox for tool orchestration, Pi’s implementation addresses tool composition, deferred loading, structured outputs, context efficiency, and coordination across agent-side and harness-side execution. |
LOCAL ORGANIZERS What’s happening across the chaptersAAIF’s local community is approaching 100 chapters worldwide, with 40 events across 15 countries in September alone. New chapters made their debut in Colombo, Luxembourg, Singapore, Shenzhen, Bengaluru, and Dallas, with more launches and meetups already on the calendar for October. Bengaluru also held its first meetup this month. Organizer Mrugesh M. thanked everyone who came along, with the next meetup already on the way.
Check out what's coming up in the events section below. |
AAIF COMMUNITY AWS Has 16,000 APIs. Can MCP Handle It?A tiny documentation change can be enough for an agent to flag trusted AWS content as a prompt-injection attack. That failure opens into a broader question: how do you make MCP servers efficient, observable, and safe once real agents start using them? - Correlating requests across an agent session reveals when ten tool calls could become one aggregate tool, cutting latency, tokens, and LLM round trips.
- Evals need to span models, harnesses, thinking settings, and tool configurations, because narrow test matrices can miss production failures.
- A constrained, schema-validatable DSL could make multi-step agent actions easier for humans to review than generated Python.
Better MCP design comes from measuring agent behavior, then shaping tools around what the data shows. Video · Spotify · Apple |
Docker's Sandbox Kit Spec puts Agent Permissions inside the container imageAn agent update can look harmless while quietly asking for broader network, credential, or filesystem access. Docker’s Sandbox Kit Spec v3 puts those permissions inside OCI images, making an agent’s authority versioned, reviewable, and portable with the artifact itself. - Permissions are declared as typed requests for network hosts, credentials, volumes, and ports, with the host approving or rejecting each one at launch.
- Updates can be compared against previously approved permissions, so non-escalating changes proceed automatically while expanded access requires review.
- Kits work with existing OCI registries, scanners, and signing tools, while conformance suites test both Kits and runtimes.
The model gives teams a concrete way to treat agent permissions as part of the software supply chain. Read the blog |
What agents look like when the data can't leave the buildingSome of the highest-value agent use cases sit on mainframes where moving production data may be restricted by policy, contract, or regulation. That changes the architecture: instead of sending records to the model, teams can move reasoning toward the system of record. - An on-prem harness translates model plans into explicit, authorized operations against DB2, IMS, CICS, VSAM, or other mainframe-native resources.
- Live reads avoid stale replicas and extra persistent copies, but introduce latency, rate-limiting, and production-load tradeoffs.
- The harness can join model traces with host-side audit records, capturing who requested what, which tools ran, and what data crossed the boundary.
For constrained environments, agent design depends on explicit capabilities, host-native authorization, and traceability across the data boundary. Read the blog |
AGENTS.md speaks UNIX, and you should tooCoding agents get much better when the environment around them is predictable, inspectable, and designed for both humans and automation. Unix tools provide the mechanics, while files such as AGENTS.md and reusable skills tell agents where to work, what to preserve, and how to check themselves. - Text files, exit codes, structured output, and composable CLI tools give agents reliable building blocks without custom integrations.
- Project instructions can separate durable source files from generated or deployed copies, preventing changes that disappear later.
- Validation commands, typed CLIs, and Git-based review keep agent work testable and explainable.
The strongest agent workflows come from improving the shared environment, instructions, and checks rather than relying on model capability alone. Read the blog |
|
Come and connect IN-PERSON EVENTSFind your city here, or start a chapter if there isn't one yet. |
Join from anywhere VIRTUAL EVENTS |
MEME OF THE WEEK  |
Sent by Demetrios Brinkmann on behalf of Agentic AI Foundation (AAIF). The Linux Foundation, 2810 N Church St., PMB 57274, Wilmington, Delaware 19802-4447, United States |
|
|