Continuing our dive into the world of Jev… (as covered in the last edition here)
LoCoMo gives Jev-Mem its clearest test: the system reaches an overall LLM-as-a-Judge score of 0.777, versus 0.700 for the strongest baseline. Its core claim is architectural: frequent memory decisions can use a lightweight System-One controller, while System Two handles complex reasoning and final answer synthesis.
The design targets a recurring cost in long-horizon agents. Autoregressive models are often asked to type memories, infer relations, route queries, score candidates, and decide when retrieval should stop. Jev-Mem turns those bounded outputs into typed probabilities, reducing token generation and parsing on both memory construction and retrieval paths.
The authors report a concrete efficiency result alongside accuracy: memory construction takes 158 seconds, compared with 1,044 seconds for the fastest competing memory system, a 6.6× speedup. Average query latency is 0.93 seconds, 36.7% below the fastest memory-based baseline at 1.47 seconds.
A quick word from our friends at Response Brief…
Where technology meets emergency response.
Response Brief covers the technology, systems, and trends changing public safety, 911, emergency communications, and how agencies respond when every second matters.
A shared control plane for writing and reading memory
Jev-Mem maintains canonical observation nodes with four overlapping relation views: semantic, temporal, causal, and entity. Vector and lexical indexes provide entry points into the same memory space, so one observation can participate in several relational structures without being duplicated across separate stores. Jev-Mem’s shared relational plane differs from hierarchical graph memory because its semantic, temporal, causal, and entity views share canonical nodes.
On the write path, System One assigns episodic, semantic, procedural, and preference type scores, then uses deterministic retrieval to select at most K_w candidate memories. It evaluates only those pairs for relatedness, causality, episode membership, and entity equivalence, while timestamps and exact identifiers can create relations without learned inference.
The system preserves every valid observation instead of making an irreversible store-or-discard decision at ingestion. Selective relation construction, bounded candidate comparison, and periodic checks for redundancy, contradiction, and obsolescence add structure while retaining original evidence. A higher-level merge can escalate to System Two only when structured control approves it.
Adaptive retrieval replaces fixed top-k search
For each query, Jev-Mem predicts which relational views are useful, how many hops may be needed, and how important recency is. It allocates a total graph-expansion budget across active views, sets traversal depth, and begins with hybrid vector-plus-lexical anchors fused by reciprocal-rank fusion.
Retrieval then cycles through route, retrieve, assess, expand, and reassess. After each round, System One estimates evidence sufficiency, expected utility of more retrieval, missing required evidence, and unresolved contradiction. Search stops when evidence is sufficient or when additional traversal is unlikely to improve the result.
Candidate scoring combines embedding similarity, query relevance, relation usefulness, novelty, evidence support, and stored edge weight, with a query-specific recency adjustment when timestamps exist. The selected memories go to System Two for synthesis; System Two does not normally control graph routing, expansion, or stopping. Jev-Mem’s explicit separation can be contrasted with System-2 memory control because Jev-Mem reserves deliberative generation for final synthesis while controlling retrieval structurally.
How Jev-Mem differs from earlier agentic memory systems
The comparison is against systems with different memory-control strategies: A-MEM dynamically evolves interconnected memory notes, Nemori uses episodic segmentation and graph retrieval, MemoryOS uses hierarchical memory tiers, and MAGMA uses semantic, temporal, causal, and entity relations. Jev-Mem differs by placing one System-One controller across both construction and retrieval rather than relying on separate heuristics or repeated generative calls. Among the named comparisons, Jev-Mem is evaluated alongside A-MEM, which dynamically evolves interconnected memory notes, but Jev-Mem applies one controller across writing and retrieval. The comparison also differs from systems that use Thompson Sampling to choose memory-management strategies, because Jev-Mem uses typed System-One decisions.
The strongest gains appear where retrieval must combine evidence or reject distractors. Jev-Mem scores 0.625 on Multi-Hop questions versus 0.569 for the strongest baseline, 0.610 on Open-Domain versus 0.517, and 0.962 on Adversarial versus 0.742. It leads in five of six LoCoMo categories and matches the best temporal score of 0.650.
The authors acknowledge a specific design tradeoff: preserving every observation avoids premature information loss, but it means the system does not discard unimportant memories at ingestion. Selectivity is deferred to relation construction and retrieval, while bounded graph, node, edge, controller-call, and latency limits prevent control overhead from growing without bound.
What the results establish about System-One memory control
Jev-Mem’s result is strongest as a systems claim supported by paired quality and runtime measurements. On LoCoMo, the same architecture that reaches 0.777 overall also reduces construction time to 158 seconds and query latency to 0.93 seconds, suggesting that lightweight control can improve evidence selection without reserving every memory decision for generation.
The reported mechanism is specific: batched typed decisions handle memory typing and relation judgments, adaptive routing allocates graph effort, candidate scoring ranks evidence, and stopping checks whether more search is worthwhile. System Two remains available for open-ended synthesis, so the architecture separates control cost from the reasoning cost of producing an answer.
The evaluation text specifies LoCoMo and the listed baselines, but it does not provide a broader cross-benchmark result in the supplied paper text. The evidence therefore supports Jev-Mem’s reported LoCoMo accuracy and efficiency claims directly, while its behavior outside that benchmark is not quantified here. The supplied evaluation reports LoCoMo results, whereas cross-scenario generality is not measured in the reported experiments.


