Skip to content

bug(memory): AND condition silently overwrites a sibling top-level filter key #7529

Description

@OfficialAbhinavSingh

Component

Python SDK

Description

Summary

In _process_metadata_filters, an AND condition whose key duplicates a sibling top-level key silently replaces that key instead of being ANDed with it. The surviving value depends on dict insertion order, so the same logical filter returns different results depending on how the caller happened to build the dict.

merge_filters deep-merges when both values are dicts, but falls through to a plain assignment for scalars (mem0/memory/main.py:1579, and the AsyncMemory twin at :3267):

if key in target and isinstance(target[key], dict) and isinstance(value, dict):
    target[key].update(value)   # operator dicts: merged  (fixed in #4853)
else:
    target[key] = value         # scalars: silently overwritten

This is the unfinished half of #4853, which fixed the operator-dict collision ({"price": {"gt": 10}} + {"price": {"lt": 20}}) and introduced this helper. The scalar collision was left overwriting.

Steps to Reproduce

No API keys needed — drives Memory against an in-memory Qdrant with a stub embedder:

from unittest.mock import patch
from mem0.memory.main import Memory
from mem0.vector_stores.qdrant import Qdrant

store = Qdrant(collection_name="repro", embedding_model_dims=3, path=":memory:")
store._bm25_encoder = False
store.insert(
    vectors=[[1.0, 0.0, 0.0]] * 2,
    payloads=[{"data": "alice-memory", "user_id": "alice", "category": "work"},
              {"data": "bob-memory", "user_id": "bob", "category": "work"}],
    ids=[1, 2],
)

class StubEmbedder:
    def embed(self, text, memory_action):
        return [1.0, 0.0, 0.0]

m = Memory.__new__(Memory)
m.vector_store, m.embedding_model = store, StubEmbedder()
m.api_version, m.reranker = "v1.1", None

def search(filters):
    with (patch("mem0.memory.main.capture_event"),
          patch("mem0.memory.main.lemmatize_for_bm25", side_effect=lambda t: t),
          patch("mem0.memory.main.extract_entities", return_value=[]),
          patch("mem0.memory.main.display_first_run_notice")):
        return sorted({r["user_id"] for r in m.search("memory", filters=filters)["results"]})

print("user_id=alice                      ->", search({"user_id": "alice"}))
print("alice AND user_id=bob              ->", search({"user_id": "alice", "AND": [{"user_id": "bob"}]}))
print("same filter, keys inserted swapped ->", search({"AND": [{"user_id": "bob"}], "user_id": "alice"}))
print("alice OR  user_id=bob              ->", search({"user_id": "alice", "OR": [{"user_id": "bob"}]}))
print("alice AND category=work (no clash) ->", search({"user_id": "alice", "AND": [{"category": "work"}]}))

Expected Behavior

{"user_id": "alice", "AND": [{"user_id": "bob"}]} means user_id == alice and user_id == bob, which nothing can satisfy, so it should return [].

docs/open-source/features/metadata-filtering.mdx states this directly:

Sibling top-level keys are implicitly ANDed on both self-hosted and the hosted Platform API, so a flat filter like {"user_id": "alice", "category": "work"} works without wrapping it in AND on either.

And the operator table lists AND as "Combine filters".

Actual Behavior

The top-level user_id is dropped and only the AND condition survives — and swapping the insertion order of the two keys flips the answer:

user_id=alice                      -> ['alice']
alice AND user_id=bob              -> ['bob']      # expected []
same filter, keys inserted swapped -> ['alice']    # expected []
alice OR  user_id=bob              -> []           # correct
alice AND category=work (no clash) -> ['alice']    # correct

Environment

  • mem0 version: 2.2.1; also present on main at abb81c88
  • Python version: 3.12
  • Vector store: Qdrant (in-memory); the defect is in mem0/memory/main.py, before any store-specific translation
  • OS: Linux

How You Verified This

What I Ran

The script above, against a clean checkout, with no LLM and no embedding API key.

What I Saw

The five lines quoted under Actual Behavior.

Why This Is a Bug

Three independent signals, not just the doc wording:

  1. It is order-dependent. The same logical filter returns ['bob'] or ['alice'] purely based on which key was inserted first. Nothing in the documented grammar makes filter results depend on dict construction order.
  2. OR gets the identical case right. {"user_id": "alice", "OR": [{"user_id": "bob"}]} correctly returns [], because OR is written to $or as a separate key and therefore still ANDs with the sibling. Only AND merges into the same dict and clobbers. That inconsistency inside one function is hard to read as intentional.
  3. It is the unfinished half of an accepted fix. fix: merge same-key operator dicts in AND metadata filters #4853 (merged) fixed the same-key collision for operator dicts and added merge_filters; scalars still take the overwrite branch.

What I Ruled Out

Suggested Fix

Make the scalar branch combine rather than replace, so a genuine contradiction yields no results instead of silently picking a winner — for example by folding a colliding scalar into an eq operator dict alongside the existing value, or by rejecting contradictory equality constraints outright. Both Memory and AsyncMemory copies would need it.

AI Assistance

AI helped me find it, and I reproduced it myself afterwards

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions