FrameworkTesting and debugging

letta-evals

Evaluates stateful agents with datasets, graders, rewards, multi-turn cases, and repeatable suites.

Stars
77
Forks
12
License
Apache-2.0
Updated
Updated Jul 8, 2026

Checking live repository facts…

Project overview

Letta Evals organizes agent evaluation from dataset and target through extractors, graders, rewards, and stored results. It supports JSONL or CSV data, multi-turn samples, multiple model handles, cached re-grading, deterministic or model-judge graders, and custom agent factories. Runs target Letta Code through a self-hosted or cloud server, so Python 3.11+, a running Letta service, and provider credentials are required.

Repository facts

Primary language
Python
License
Apache-2.0
Repository updated
Jul 8, 2026
Default branch
main

Resource types

FrameworkBenchmark or evalCLI app

Use cases

Testing and debuggingAI and agent developmentData and analytics

Runtime

Command lineCloudLocalSandboxedSelf-hosted

Protocols & integrations

CLI integration

Capabilities

Code executionData retrievalObservabilityVerification and evalsWorkflow automation

Audience

DevelopersResearchers

Related projects

Browse more similar projects