Generating Long Sequences with Sparse Transformers
TM
Trevor McFedries
@trevvyboi
Transformers are powerful sequence models, but require time and memory that grows quadratically with the sequence length. In this paper we introduce sparse factorizations of the attention matrix which reduce this to $O(n \sqrt{n})$. We also introduce a) a v...
- Uploaded
- Uploaded Jul 10, 2026
- Queried
- Queried 0 times
No preview text is available for this document yet.
Want to learn more?
Ask a question