The worry now is people in the EU mistake my typos for a watermark. |
In case you missed it GemsMCP Associate exam prep guide Built around the official exam blueprint, this free 34-lesson path covers MCP architecture, JSON-RPC, tools, manifests, security boundaries, diagnostics, mock tests, and study plans for independent preparation. Agent memory systems benchmark A reproducible harness compares Mem0, Letta, LangMem, Zep/Graphiti, and GoodMem on retrieval quality, adversarial refusal, latency, context size, and ingestion time using the same long-conversation workload. Large tool response offloading guide Covering threshold handling, sandbox offloading, progressive inspection, parallel-response limits, and security controls, it gives practitioners a practical way to reduce context pressure without discarding full tool payloads. Self-hosted agent sandbox tool Built for isolated code execution, it gives agents disposable Docker sandboxes with snapshots, rollback, forking, TTL cleanup, and access through a REST API, Python SDK, or MCP server. |
FREE VIRTUAL EVENT Accuracy & Reliability of AI Agents What does “reliable” mean for an AI agent, and how do we measure it?
This virtual summit will look at how teams are evaluating agent behavior in practice, where current benchmarks fall short, and what we’re learning from failures in real-world systems. October 15 | 8:00-10:00 AM PT / 3:00-5:00 PM UTC JOIN LIVE |
AAIF COMMUNITY The caveman prompting challengeA $10 AI task only becomes meaningful when you can say what that $10 produced. The focus here is unit economics for AI workloads: tying token spend to business outcomes, then optimizing cost, speed, and accuracy around a measurable target. - Start with the strongest model, then step down once the task works - for example, from a 300B-parameter model to 70B with more orchestration around it.
- Treat prompt efficiency, model choice, latency, and retries as connected cost controls rather than isolated tweaks.
- Keep agent governance close to existing cloud practice: policy as code, automated guardrails, APIs, and production pipelines.
The useful unit is cost per outcome, not tokens consumed in isolation. Video · Spotify · Apple |
How a Logistics Giant Keeps AI Data Locked DownA runaway retry loop can turn prompt bloat into a real cost problem, with every failed pass consuming the same tokens again. The broader challenge is measuring AI workloads by value and efficiency, not spend alone. - Tag AI apps and agentic workflows by owner, then track cost and efficiency by app or model.
- Separate conversational and agentic workloads before judging context size; multi-step agents can look bloated when several prompts are treated as one.
- Watch cache hit rate, reasoning-token use, retries, and context starvation; prompt caching alone can cut costs by 60–70%.
Good observability makes AI cost signals useful for debugging, architecture choices, and ROI measurement. Video · Spotify · Apple |
Trace and analyze Goose sessions with OpenTelemetry, Jaeger, and ClickHouseA three-word greeting took Goose 14 seconds to answer, but tracing showed the model was responsible for only about half of that delay. Adding OpenTelemetry, Jaeger, and ClickHouse exposed what was happening across model calls, agent processing, tool use, and stored traces. - Switching free-tier models cut the provider call from ~7 seconds to ~1 second, while ~7 seconds of agent-side latency remained.
- A Cloudflare MCP task took 48.9 seconds across 10 alternating model and tool-call spans.
- ClickHouse kept Jaeger traces available across container restarts for later inspection.
Tracing separates model latency from orchestration overhead, making slow agent behavior much easier to diagnose. Read the blog |
There is no one agentic commerce protocolAgentic commerce gets complicated the moment an agent moves from finding a product to spending someone else’s money. The emerging stack splits that journey across discovery, carts, delegated authority, payments, coordination, fulfillment, and returns. - UCP exposes commerce capabilities such as discovery and portable carts, while AP2 adds cryptographic proof of what a user authorized.
- x402 handles web-native payments, A2A coordinates agents across systems, and neither replaces the surrounding commerce flow.
- Real purchase journeys still require explicit handoffs across protocols, payment rails, merchant systems, tracking, and returns.
The useful architectural boundary is between what an agent may read freely and where verified authority to spend must begin. Read the blog |
Agent identity and delegated access in MCP systemsOne GitHub update can cross several agents, MCP servers, and downstream services, each making its own identity and authorization decision. The challenge is preserving who initiated the work, who is acting now, and exactly which permissions should survive each hop. - User-delegated OAuth, workload identities, and RFC 8693 token exchange cover different authorization patterns without collapsing everything into one shared credential.
- Delegated agents should receive narrower, task-specific authority, with reauthorization where long-running tasks outlive approvals or permissions.
- OpenTelemetry, W3C Trace Context, and MCP trace propagation help reconstruct execution alongside authorization records.
Reliable agent access depends on keeping identity, delegated authority, and execution history connected across the full request path. Read the blog |
|
LOCAL ORGANIZERS What’s happening across the chaptersAustralia had a busy week, with Melbourne attendees hearing about Goose, ACP, and what AI-assisted development could mean for junior developers, while Sydney held its first AAIF community event. Sydney’s talks covered maturing MCP, multi-cloud X-ops agents, and measuring user value for MCP servers. Ahmedabad also held its first meetup, covering multi-agent systems and observability alongside custom harnesses and model routing. Seattle brought more than 50 developers together for durable agents with Dapr Agents and an introduction to A2A. In London, 687 people registered for an evening spanning post-training, RL, evals, inference, and agent architecture, with around 64% of the approved list turning up. |
Come and connect IN-PERSON EVENTSFind your city here, or start a chapter if there isn't one yet. |
Join from anywhere VIRTUAL EVENTS |
MEME OF THE WEEK  |
Sent by Demetrios Brinkmann on behalf of Agentic AI Foundation (AAIF). The Linux Foundation, 2810 N Church St., PMB 57274, Wilmington, Delaware 19802-4447, United States |
|
|