Frontline Lab summary and source
Agent memory is not a feature to switch on casually; its dose must be calibrated to model capability. Strong models are better suited to injecting a full set of guides, with DeepSeek-V3.2 (671B MoE) improving task completion by +9.5 percentage points. Weaker models perform best with curated retrieval, with gpt-oss-120b (117B MoE) improving by +16.1pp while adding only +5% tokens. This method requires no weight updates or manual annotation; it works by distilling guides from an agent’s past trajectories and injecting them at inference time.
This brief preserves the original source so the summary and editorial context can be checked independently.
Source attributionHugging Face:Blog(RSS) · Hugging Face:Blog(RSS)
Open the original source