Frontline Lab summary and source
@dair_ai said Ai2, Carnegie Mellon, and University of Washington studied normalization, GQA, pretraining context length, and sliding-window attention, finding that combining more than three of them significantly reduces long-context performance, and released OlmPool.
This brief preserves the original source so the summary and editorial context can be checked independently.
Source attributionX · @dair_ai
Open the original source