Component
Python SDK
Description
I've been building a local faithfulness-checking library (MiniEval - (https://github.com/DataAlchmesit/MiniEval-pro), MIT licensed, 'pip install minieval-pro' and ran into a failure mode I think is worth raising here.
The problem: When an NLI-style faithfulness model checks a candidate fact against its source, it has no notion of whose fact it is. Given a source like "my brother is a lawyer" and a candidate "the user is a lawyer," a faithfulness model can score that as ~99% entailed. The professions match, so it entails, even though the fact is misattributed. I've tested this specifically and built two guards against it (possessive/reported-speech detection, and a relatedness check for cases where an unrelated pair gets forced into a false "contradiction" by the NLI model).
I looked through mem0/memory/main.py to understand where this would actually apply, and found the extraction happens in '_add_to_vector_store', a single LLM call parses extracted_memories from the conversation (around the "V3 PHASED BATCH PIPELINE" section), which then go through hashing/dedup and batch insert. That gap between "extraction parsed" and "batch persist" looks like the natural place a verification step could sit, but I want to ask rather than assume:
Is there already a supported way to hook into or intercept the extraction result before it's persisted , a callback, a config option, an extension point I'm missing?
If not, would a maintainer be open to discussing what a minimal, opt-in verification hook might look like? I'd rather design it with your input than build something against a private method (_add_to_vector_store) that could break on your next refactor.
To be clear about scope: I'm not proposing MiniEval become a dependency or that Mem0 adopt any particular faithfulness model. I'm asking whether an extension point for this class of check (verify before persist) is something you'd want, and if so, what shape it should take on your side.
Happy to share the specific test cases I've built (including ones where my own approach still fails, documented honestly) if useful context for the discussion.
AI Assistance
No AI involved, this is a need I hit myself
Component
Python SDK
Description
I've been building a local faithfulness-checking library (MiniEval - (https://github.com/DataAlchmesit/MiniEval-pro), MIT licensed, 'pip install minieval-pro' and ran into a failure mode I think is worth raising here.
The problem: When an NLI-style faithfulness model checks a candidate fact against its source, it has no notion of whose fact it is. Given a source like "my brother is a lawyer" and a candidate "the user is a lawyer," a faithfulness model can score that as ~99% entailed. The professions match, so it entails, even though the fact is misattributed. I've tested this specifically and built two guards against it (possessive/reported-speech detection, and a relatedness check for cases where an unrelated pair gets forced into a false "contradiction" by the NLI model).
I looked through mem0/memory/main.py to understand where this would actually apply, and found the extraction happens in '_add_to_vector_store', a single LLM call parses extracted_memories from the conversation (around the "V3 PHASED BATCH PIPELINE" section), which then go through hashing/dedup and batch insert. That gap between "extraction parsed" and "batch persist" looks like the natural place a verification step could sit, but I want to ask rather than assume:
Is there already a supported way to hook into or intercept the extraction result before it's persisted , a callback, a config option, an extension point I'm missing?
If not, would a maintainer be open to discussing what a minimal, opt-in verification hook might look like? I'd rather design it with your input than build something against a private method (_add_to_vector_store) that could break on your next refactor.
To be clear about scope: I'm not proposing MiniEval become a dependency or that Mem0 adopt any particular faithfulness model. I'm asking whether an extension point for this class of check (verify before persist) is something you'd want, and if so, what shape it should take on your side.
Happy to share the specific test cases I've built (including ones where my own approach still fails, documented honestly) if useful context for the discussion.
AI Assistance
No AI involved, this is a need I hit myself