← Mnemos

02 / Mnemos - eval fixture

Archived: this early illustrative episode no longer represents the active Mnemos evaluation design. See the current evidence-first HippoCamp plan.

Long-history memory episode

One inspectable sample: six independent OpenHarness conversations become a single chronological memory stream, then three fresh-session questions test write, retrieval, and final-answer quality.

Download raw JSON
Harness
OpenHarness
Model
Loading...
Source conversations
6
Raw events
34
History window
19 days
Built-in / MCP calls
2 / 6
Evaluation questions
3
Fixture schema
mnemos.eval.fixture.v1

System under evaluation

OpenHarness sessions6 separate conversations
Raw trajectoriesmessages + tool IO
Normalizeridentity, time, source IDs
Memory writeextract, reject, supersede
Memory retrievalbackend-specific context
Agent + gradersanswer and 3 score layers
Source conversations OpenHarness executions grouped by session

Loading fixture...

Complete seed trace 34 normalized events in ingestion order
#TimestampConversationKindRole / toolMemory signal
Memory pipeline run normalized input, write candidates, and rejected noise

Loading fixture...

Question runs retrieval context, final answer, and layer-specific graders

Loading fixture...