Skip to content

Explore Jev as an optional semantic judgment layer for memory retrieval #1703

Description

@romanserk

Idea

Would it be worth benchmarking Jev as an optional listwise semantic judgment layer in Supermemory retrieval?

Supermemory already supports cross-encoder reranking, so the interesting question may not be “can Jev rerank?”, but whether its typed, listwise judgments are useful after candidate generation:

hybrid / graph retrieval
        ↓
candidate set
        ↓
existing reranker and/or Jev listwise judgment
        ↓
final context

A useful reference point is Hindsight's recent Jev integration. Instead of judging each candidate independently, Hindsight turns the candidate pool into the options of a single Choice question and uses the returned probability distribution as the ranking.

Their implementation notes report that, on a 200-question LoCoMo set, the listwise approach reached Recall@1 0.94 vs 0.87 for one-call-per-candidate scoring, while using about 1/30th the calls.

They also published a direct comparison against their MiniLM cross-encoder:

  • 30 candidates: 0.950 vs 0.800 Recall@1
  • 240 candidates: 0.783 vs 0.583 Recall@1

Those numbers are specific to Hindsight, so I would treat them as motivation to benchmark rather than an expectation for Supermemory.

One Supermemory-specific angle that seems worth testing is whether Jev benefits from the context Supermemory already has around a memory — for example current-vs-historical state and related memories — rather than judging only the raw memory text.

I would probably keep dynamic pruning as a separate experiment. Hindsight supports a second Score question to choose where relevance ends, but keeps it opt-in; in their published 30-candidate test it removed 19% of gold evidence. That makes it look more like a context-size / precision tradeoff than an unconditional retrieval improvement.

Possible benchmark

  • current cross-encoder reranking
  • Jev listwise reranking on the same candidate set
  • current reranking + Jev as a final judgment layer

Metrics could include Recall@K / NDCG, final context size, latency, and cost.

References:


On a personal note, I really like what you’re building with Supermemory and have been using it quite a bit recently.

I wrote about my experience using Supermemory as a shared memory layer across AI agents:

So this suggestion comes from using the product and thinking about where this kind of decision model might fit naturally.

Would really appreciate your thoughts on whether this is a direction worth exploring.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions