Idea
Would it be worth benchmarking Jev as an optional listwise semantic judgment layer in Supermemory retrieval?
Supermemory already supports cross-encoder reranking, so the interesting question may not be “can Jev rerank?”, but whether its typed, listwise judgments are useful after candidate generation:
hybrid / graph retrieval
↓
candidate set
↓
existing reranker and/or Jev listwise judgment
↓
final context
A useful reference point is Hindsight's recent Jev integration. Instead of judging each candidate independently, Hindsight turns the candidate pool into the options of a single Choice question and uses the returned probability distribution as the ranking.
Their implementation notes report that, on a 200-question LoCoMo set, the listwise approach reached Recall@1 0.94 vs 0.87 for one-call-per-candidate scoring, while using about 1/30th the calls.
They also published a direct comparison against their MiniLM cross-encoder:
- 30 candidates: 0.950 vs 0.800 Recall@1
- 240 candidates: 0.783 vs 0.583 Recall@1
Those numbers are specific to Hindsight, so I would treat them as motivation to benchmark rather than an expectation for Supermemory.
One Supermemory-specific angle that seems worth testing is whether Jev benefits from the context Supermemory already has around a memory — for example current-vs-historical state and related memories — rather than judging only the raw memory text.
I would probably keep dynamic pruning as a separate experiment. Hindsight supports a second Score question to choose where relevance ends, but keeps it opt-in; in their published 30-candidate test it removed 19% of gold evidence. That makes it look more like a context-size / precision tradeoff than an unconditional retrieval improvement.
Possible benchmark
- current cross-encoder reranking
- Jev listwise reranking on the same candidate set
- current reranking + Jev as a final judgment layer
Metrics could include Recall@K / NDCG, final context size, latency, and cost.
References:
On a personal note, I really like what you’re building with Supermemory and have been using it quite a bit recently.
I wrote about my experience using Supermemory as a shared memory layer across AI agents:
So this suggestion comes from using the product and thinking about where this kind of decision model might fit naturally.
Would really appreciate your thoughts on whether this is a direction worth exploring.
Idea
Would it be worth benchmarking Jev as an optional listwise semantic judgment layer in Supermemory retrieval?
Supermemory already supports cross-encoder reranking, so the interesting question may not be “can Jev rerank?”, but whether its typed, listwise judgments are useful after candidate generation:
A useful reference point is Hindsight's recent Jev integration. Instead of judging each candidate independently, Hindsight turns the candidate pool into the options of a single
Choicequestion and uses the returned probability distribution as the ranking.Their implementation notes report that, on a 200-question LoCoMo set, the listwise approach reached Recall@1 0.94 vs 0.87 for one-call-per-candidate scoring, while using about 1/30th the calls.
They also published a direct comparison against their MiniLM cross-encoder:
Those numbers are specific to Hindsight, so I would treat them as motivation to benchmark rather than an expectation for Supermemory.
One Supermemory-specific angle that seems worth testing is whether Jev benefits from the context Supermemory already has around a memory — for example current-vs-historical state and related memories — rather than judging only the raw memory text.
I would probably keep dynamic pruning as a separate experiment. Hindsight supports a second
Scorequestion to choose where relevance ends, but keeps it opt-in; in their published 30-candidate test it removed 19% of gold evidence. That makes it look more like a context-size / precision tradeoff than an unconditional retrieval improvement.Possible benchmark
Metrics could include Recall@K / NDCG, final context size, latency, and cost.
References:
On a personal note, I really like what you’re building with Supermemory and have been using it quite a bit recently.
I wrote about my experience using Supermemory as a shared memory layer across AI agents:
So this suggestion comes from using the product and thinking about where this kind of decision model might fit naturally.
Would really appreciate your thoughts on whether this is a direction worth exploring.