02 / Mnemos - eval fixture
Archived: this early illustrative episode no longer represents the active Mnemos evaluation design. See the current evidence-first HippoCamp plan.
Long-history memory episode
One inspectable sample: six independent OpenHarness conversations become a single chronological memory stream, then three fresh-session questions test write, retrieval, and final-answer quality.
System under evaluation
OpenHarness sessions6 separate conversations
Raw trajectoriesmessages + tool IO
Normalizeridentity, time, source IDs
Memory writeextract, reject, supersede
Memory retrievalbackend-specific context
Agent + gradersanswer and 3 score layers
Source conversations OpenHarness executions grouped by session
Loading fixture...
Complete seed trace 34 normalized events in ingestion order
| # | Timestamp | Conversation | Kind | Role / tool | Memory signal |
|---|
Memory pipeline run normalized input, write candidates, and rejected noise
Loading fixture...
Question runs retrieval context, final answer, and layer-specific graders
Loading fixture...