Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

📡 Tech & Science

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

arXiv:2609.20888v1 Announce Type: new Abstract: Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense model quality. ETA predicts dynamic, contextual thresholds directly from query representations, allowing the model to allocate dense-like context to difficult retrieval or reasoning steps while pruning routine tokens. To learn this policy from scratch without representation collapse, ETA \emph{multiplicatively suppresses} sub-threshold logits toward zero during training rather than deleting them. Training against this smooth uniform attention floor provides a distributed probability reservoir that \textbf{causes localized attention sinks on initial tokens to disappear}. It also enables the model to hard-prune uninformative KV blocks at inference time and absorb incidental tokens co-admitted by coarse GPU block selection. As a result, a 1.45B pretrained ETA model rivals dense attention across language modeling, commonsense reasoning, and long-context needle retrieval at $\approx 85\%$ training sparsity and $\approx 38\%$ active decode density. At inference time, we implement a custom decode kernel in Triton that screens KV blocks in $O(1)$ time using cached geometric-probabilistic bounds, delivering up to $2.5\times$ wall-clock decode speedups over FlashAttention-2 on sequences up to 512K tokens. Finally, we introduce an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds to eliminate predictor overhead, cutting attention compute by an additional $27\%$.

📖 Cet article provient d'une source externe.

🔗 Lire l'article complet sur la source →

246 mots extraits · Source originale


🔥 OFFRE PARTENAIRE

Xiaomi POCO M7 Global, 6 Go + 128 Go, lecteur d'empreintes digitales latéral, écran 6,9 pouces, Xiaomi HyperOS 2, Snapdragon 685 4G Octa Core, NFC, réseau 4G (bleu)

🔥 Xiaomi POCO M7 Global, 6 Go + 128 Go, lecteur d'empreintes digitales latéral, écran 6,9 pouces, Xiaomi HyperOS 2, Snapdragon 685 4G Octa Core, NFC, réseau 4G (bleu) - Une offre exceptionnelle à ne pas manquer ! Cliquez pour découvrir.
✅ Consultez les photos supplémentaires.

✅ Découvrez toutes les caractéristiques.

✅ Vérifiez la disponibilité actuelle.

✅ Consultez les avis des acheteurs.

Posts les plus consultés de ce blog

Comment mettre un accent à une lettre majuscule À, É, È, Ç, Î, Ô, Û pour Windows

Voici 50 raccourcis clavier utiles dans Microsoft Word (version Windows)

Gigantic Jet 

Better tools made Copilot code review worse. Here’s how we actually improved it.

Monstre d’acier : avec une hauteur de 250 mètres, Big Carl est la plus grande grue du monde

Geoffrey Hinton, le prix Nobel de physique qui a démissionné de Google et dénoncé les dangers de l'intelligence artificielle pour l'humanité

Show HN: CheapSecurity – Lightweight, Self-Hosted CCTV for Linux SBCs

Comment copier-coller le texte d’une image ?Extraire et copier le texte d'une image avec l'outil Capture d'écran

⚡ Weekly Recap: Rogue AI Agents, Check Point Exploit, Slopsquatting, ClickFix Lures and More