Skip to content

extract_json silently drops memories when an LLM wraps JSON in prose containing braces #5998

Description

@ritsth

What happened

mem0.memory.utils.extract_json is the last-resort parser for LLM responses with no code fence. It returns the span from the first { to the last }. When a chatty LLM (common with local Ollama / LM Studio models) wraps its JSON in prose that itself contains braces, that span begins inside the prose and is not valid JSON. json.loads raises, the outer except in _add_to_vector_store swallows it, and the extracted memories are silently dropped with no error surfaced to the caller.

This is the same silent-memory-loss class as #5245 and #5509, and it falls in the scope tests/test_chatty_llm_parsing.py already targets, but the existing tests only cover prose without braces.

Steps to reproduce (offline, no API key)

import json
from mem0.memory.utils import extract_json, remove_code_blocks

response = 'Based on the conversation {about travel}, here is the update: {"memory": [{"id": "0", "text": "likes travel", "event": "ADD"}]}'

# Mirrors the fallback chain in _add_to_vector_store
text = remove_code_blocks(response)
try:
    parsed = json.loads(text, strict=False)
except json.JSONDecodeError:
    parsed = json.loads(extract_json(text), strict=False)  # raises -> caught upstream -> memories dropped

extract_json returns '{about travel}, here is the update: {"memory": [...]}', which fails to parse, so the memory is lost with no error.

Expected behavior

The JSON object ({"memory": [...]}) should be extracted and stored, regardless of brace-containing prose around it.

Environment

mem0 main (reproduced on 2.0.10), Python 3.11, any LLM that returns conversational text around the JSON.

Activity

  1. iroiro147 commented on Jul 8, 2026

    @iroiro147

    Hi, I can take this. I’ll add a regression for brace-containing prose around the JSON payload and update the fallback extractor so it returns the valid JSON object instead of the first-brace/last-brace span.

  2. hegu-1 commented on Jul 12, 2026

    @hegu-1

    Silent memory loss is especially dangerous because it looks like a model/personality problem later: the agent simply "forgot," but the actual failure happened at extraction time.

    For memory systems, I would treat parser failure as a first-class audit event rather than only an exception path. Even if the JSON cannot be recovered, the system should preserve enough metadata to explain what happened:

    memory_write_attempted
    parser: extract_json
    status: failed | recovered
    raw_response_hash
    error_class
    fallback_used
    

    In a personal memory vault pattern, I keep raw events separate from derived memory exactly for this reason: if a derived node fails to build, the source event still exists and can be reprocessed later.

    Reference shape:
    https://github.com/hegu-1/personal-memory-vault-starter

    For Mem0, a robust extractor fix is the direct issue, but surfacing "memory extraction failed" to the caller or telemetry would also help. Otherwise local/chatty models can silently create holes in long-term continuity that are very hard to debug afterwards.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions