New Tools from Diagrid and AWS Target AI Agent Durability and Evaluation
Recent infrastructure releases from Diagrid and AWS are introducing standardized durability, verifiable execution, and framework-agnostic evaluation for multi-agent workflows.

Key takeaways · 3
- 01
Diagrid Catalyst 2.0 enables interrupted agent runs to resume without repeating completed model calls.
- 02
Amazon Bedrock AgentCore leverages OpenTelemetry to decouple agent evaluation from specific framework choices.
- 03
Catalyst's cryptographic history signing relies on Dapr 1.18 and requires mTLS to be enabled.
Adding Durability and Verification
Diagrid announced Catalyst 2.0 on July 28, 2026, adding failure recovery and cryptographic verification to agents built with frameworks like LangGraph and Dapr Agents. [1] Catalyst represents model and tool calls as durable workflow activities, which allows interrupted runs to resume without repeating already completed work. [1]
The verification model, stemming from Dapr 1.18, hashes batches of workflow-history events, links each digest to the preceding signature, and signs the result to make deleted, reordered, or modified history detectable. [1] This signing feature is disabled by default and depends on mTLS being enabled. [1]
Framework-Agnostic Evaluation
Amazon Bedrock AgentCore evaluations solves evaluation fragmentation by decoupling the evaluation process from the specific agent framework choice. [2] The service can score an agent regardless of the underlying SDK as long as its telemetry flows through OpenTelemetry. [2] This capability supports developers building on frameworks like LangGraph, Google ADK, and Strands Agents deployed on the Amazon Bedrock AgentCore runtime. [2]
What it means
The AI agent ecosystem is shifting from isolated framework silos to standardized runtime and operational tooling. By standardizing on Dapr 1.18 for identity and OpenTelemetry for telemetry, vendors are addressing the operational fragility of long-running multi-agent workflows. Compared to earlier bespoke evaluation setups, Amazon's approach allows centralized observability across disparate SDKs, while Diagrid tackles the high compute costs associated with late-stage execution failures in multi-step runs. What the sources don't address: How these operational overheads, such as cryptographic history signing and telemetry emission, impact overall agent response latency in production environments.
As enterprise AI teams deploy complex multi-agent workflows, managing failure recovery and standardizing evaluation have become critical hurdles. These tools provide infrastructure-level solutions to ensure robust execution and consistent performance metrics across varied architectures.
Why it matters
Turn this story into practical AI skill after launch.
Get the release link for daily sessions built around your role and industry.
Join the waitlistHow this developed
28 August 2026
New Tools from Diagrid and AWS Target AI Agent Durability and Evaluation
28 August 2026
Event created from source cluster.