Published collection
OlmPool shows four architecture choices weaken long context
Research suggests short-context validation may miss long-context degradation, affecting model architecture trade-offs.
- Published entries
- 1
- Sources
- 1
- Date range
- 2026-08-12
- Primary labels
- Research · 论文 · Research Papers · OlmPool 长上下文架构研究 · Meta · 阿里千问
Published evidence
Every entry keeps its summary and a path back to the source context.
OlmPool shows four architecture choices weaken long context
@dair_ai said Ai2, Carnegie Mellon, and University of Washington studied normalization, GQA, pretraining context length, and sliding-window attention, finding that combining more than three of them significantly reduces long-context performance, and released OlmPool.
Original source