implementation detail · filed under model architecture
IndexCache
A mechanism for cross-layer sparse-index reuse.
- sources
- 3
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 3
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse paper arxiv.orgintroduces IndexCache cross-layer sparse-index reuse
Evidence
3 spans quoted from the sources, strongest treatment first.
with IndexCache for cross-layer sparse index reuse
coremodel architecturein Hy4-previewTencent Hunyuan
with IndexCache for cross-layer sparse index reuse
coremodel architecturein Hy4-previewTencent Hunyuan
with IndexCache for cross-layer sparse index reuse
coremodel architecturein Hy4-previewTencent Hunyuan
Filed alongside
Other methods under model architecture :: token mixer :: sparse attention :: sparse attention indexer.
Compressed Sparse Attention 2Hierarchical Sparse IndexerIndexShareLightning IndexerReindex ModeReuse ModeDense Warm-up StageFine-grained token selectionIndex BranchIndexer WarmupReuse QSA index selection across speculative decoding stepsAverage poolingBlock-causal scoringCompressed lightweight indexerCross-stage shared-state management for attention reuseDetached indexer-input optimizationIndex Branch outputIndex Branch value headIndexPoolMQA indexerSingle-head index key