Publications
-
1
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context DecodingWe design a novel attention mechanism that allows the model to self-predict sparsity patterns at inference-time, thus bridging the gap between training-time quality and inference speed.
-
2
Testing $k$-submodularity
-
3
Efficient Algorithms for Influence Maximization in General Models and Observed Cascades
-
4
Layerwise Dynamics for In-Context Classification in TransformersWe show how Transformers geometrically separate classes layer-by-layer during in-context learning.
-
5
Noise Stability of Transformer ModelsWe introduce noise stability to measure model simplicity and accelerate Transformer training and grokking.
-
6
Fast-MWEM: Private Data Release in Sublinear Time
-
7
Efficient Algorithms for Adversarially Robust Approximate Nearest Neighbor Search
- NeurIPS 2025 Workshop: Reliable ML from Unreliable Data
- WoLA 2026 Poster
-
8
Estimating Hitting Times Locally At Scale
-
9
Compression Barriers for Autoregressive TransformersWe prove fundamental limits on KV cache compression, showing when sublinear memory is impossible.
-
10
$k$NN Attention Demystified: A Theoretical Exploration for Scalable TransformersWe provide the first theoretical guarantees and fast sub-quadratic algorithms for $k$NN attention.
-
11
Counting Simplices in Hypergraph Streams
-
12
Teaching American Sign Language in Mixed Reality