About Me
I am Mohammad Mozaffari, an ML Researcher at ElastixAI. I received my PhD in Computer Science from the University of Toronto, supervised by Professor Maryam Mehri Dehnavi. I got my B.Sc. in Electrical Engineering with a minor degree in Computer Engineering from the University of Tehran.
My research focuses on the "Compression Trinity" for Large Language Models: the interplay of sparsity, quantization, and low-rank approximations to make LLMs faster and smaller. My work has been featured by NVIDIA Research and the official PyTorch blog. You can explore it here: The Compression Trinity for LLMs, or read the full treatment in my Ph.D. thesis: Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
Publications
PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs
@inproceedings{hourri2025patch,
title = {{PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs}},
author = {Hourri, Younes and Mozaffari, Mohammad and Mehri Dehnavi, Maryam},
year = 2026,
booktitle = {The Fortieth Conference on Neural Information Processing Systems}
}
- Introduced hybrid dense + 2:4 sparsity with learnable tile-level masks for LLMs.
- Achieved 1.18×–1.38× inference speedup on LLaMA-2 7B with up to +2.96% accuracy vs. MaskLLM.
- Integrated with STOICC compiler to enable GPU-accelerated hybrid sparsity.
OPTIMA: Optimal One-Shot Pruning for LLMs via Quadratic Programming Reconstruction
@inproceedings{mozaffari2025optima,
title = {{OPTIMA: Optimal One-Shot Pruning for LLMs via Quadratic Programming Reconstruction}},
author = {Mozaffari, Mohammad and Kushnir, Samuel and Mehri Dehnavi, Maryam and Yazdanbakhsh, Amir},
year = 2026,
booktitle = {The Eighteenth Asian Conference on Machine Learning}
}
- Reformulated weight reconstruction as parallel row-wise QPs with shared Hessians.
- Developed accelerator-friendly QP solver, pruning an 8B LLaMA end-to-end in 40h on a single H100.
- Improved average accuracy by up to 2.53% absolute across multiple LLM families and sparsity regimes.
SLiM: One-shot Quantized Sparse Plus Low-rank Approximation of LLMs
@article{mozaffari2025slim,
title = {{SLiM: One-shot Quantized Sparse Plus Low-rank Approximation of LLMs}},
author = {Mozaffari, Mohammad and Yazdanbakhsh, Amir and Mehri Dehnavi, Maryam},
year = 2025,
url = {https://openreview.net/forum?id=4UfRP8MopP}
}
- Achieved 5.8% accuracy improvement via 2:4 sparse + quantized + low-rank compression.
- SLiM featured by NVIDIA Research.
SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs
@article{mozaffari2024slope,
title = {{SLoPe: Double-Pruned Sparse Plus Lazy Low-rank Adapter Pretraining of LLMs}},
author = {Mozaffari, Mohammad and Yazdanbakhsh, Amir and Zhang, Zhao and Mehri Dehnavi, Maryam},
year = 2025,
booktitle={The Thirteenth International Conference on Learning Representations},
url={https://openreview.net/forum?id=lqHv6dxBkj}
}
- 1.25× faster training and 1.54× faster inference using double pruning + low-rank adapters.
- PyTorch team featured SLoPe in a research blog post.
- Developed CUDA kernels for N:M semi-structured sparsity and lazy LoRA.
MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates
@inproceedings{
mozaffari2023mkor,
title = {{MKOR}: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates},
author = {Mohammad Mozaffari and Sikan Li and Zhao Zhang and Maryam Mehri Dehnavi},
booktitle ={Thirty-seventh Conference on Neural Information Processing Systems},
year ={2023},
url ={https://openreview.net/forum?id=jcnvDO96N5}
}
- Proposed a momentum-based distributed optimizer using Kronecker-factored gradients with rank-1 updates, reducing per-step communication cost.
- Achieved 2.57x wall-clock training speedup on BERT-Large (64 A100 GPUs) vs. standard first-order optimizers, with matching convergence.
ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences
@article{chen2026rote,
title = {{ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences}},
author = {Chen, Xinye and G{\"u}ttel, Stefan and Mozaffari, Mohammad},
year = 2026,
journal = {Under review},
url = {https://hal.science/hal-05765229}
}
- Introduced ROTE, a benchmark that controls symbolic-sequence complexity via LZW compression and tests both teacher-forced next-symbol prediction and closed-loop rollout.
- Benchmarked 8 recurrent, attention-based, and hybrid architectures (LSTM, GRU, minGRU, minLSTM, Transformer, LinearAttention, Performer, RWKV) over 640 training runs, reporting accuracy alongside parameters, training time, and peak memory.
- Showed that model size does not predict memorization: GRU (0.15M parameters) reproduces 92.5% of 100-symbol rollouts exactly vs. 61.3% for a 3.7M-parameter Transformer.
Workshop Papers
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
@inproceedings{mozaffari2026leap,
title = {{LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models}},
author = {Mozaffari, Mohammad and Hourri, Younes and Rastegari, Mohammad and Najibi, Mahyar},
year = 2026,
booktitle = {ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)}
}
- Improved unstructured sparsity accuracy by 5.40% via per-weight learnable masks (Gumbel-sigmoid relaxation).
Blog Posts
- The Compression Trinity for LLMs — A survey of sparsity, quantization, and low-rank approximation, and how they combine Updated Sep 2026
- LEAP: Learnable End-to-End Adaptive Pruning of LLMs — Learnable per-weight masks for unstructured sparsity Dec 2025
- SLICE: Layer-wise Configuration Search for Matryoshka LLMs — Choosing per-layer width and precision under a budget Dec 2025
- BEAM: Blockwise Error Minimization for One-shot Compression of LLMs — Recovering accuracy after compression without full retraining Jun 2025
Invited Talks
- PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs — NVIDIA Research, Seattle Oct 2025
- Compression Trinity: Interplay of Sparsity, Quantization, and Low-Rank Approximation for LLMs — Cerebras, Toronto Mar 2025
- Efficient LLM Training and Inference: Sparsity, Quantization, and Low-Rank Approximation — Google DeepMind, Seattle Mar 2025
- Enabling Semi-structured Sparsity in LLMs — NVIDIA Research, Seattle Mar 2024
- Communication-Efficient Second-Order Optimization Methods — Rutgers University, New Jersey Nov 2023
Media & Outreach
- When Quantization Isn't Enough: Why 2:4 Sparsity Matters — Official PyTorch Blog
-
Guest Interview on Executive Code Podcast — Smaller Models, Same Power: How SLiM Shrinks LLMs Without Retraining Jul 2025
YouTube · Spotify · Apple Podcasts
Mentorship
Mentored 7 undergraduate and master's students on projects related to LLM compression. Two mentees were admitted to Stanford for graduate studies.
Experience
ML Researcher at ElastixAI Dec 2025 – Present
- Research and develop compression techniques for efficient deployment of large language models.
- Investigate Mixture-of-Experts architectures, including token routing, kernel design, and dispatch optimization.
- Collaborate directly with the CTO and CEO on research direction and production integration.
Research Intern at Autodesk Aug 2022 – Dec 2022
- Reduced multi-GPU simulation time from 4h to 3.2h via CUDA kernel optimization and profiling.
- Designed kernel fusion and memory coalescing strategies, reducing bandwidth by 30%.
- Improved inter-GPU synchronization and dataflow using Nsight Systems.