REVIEW 2 major objections 1 minor 3 cited by
MixLLM assigns bit widths to LLM output features by global importance rather than per layer to cut accuracy loss at modest extra memory cost.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-05-23 06:53 UTC
load-bearing objection MixLLM's global output-feature ranking for mixed-precision quantization is the claimed novelty but the abstract supplies no method details, ablations, or normalization steps, so the accuracy gains cannot be assessed yet. the 2 major comments →
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MixLLM identifies the important output features in the global view rather than within each single layer, effectively assigning larger bit-width to output features that need it the most to achieve high accuracy and low memory usage, while the two-step dequantization and overlapping software pipeline deliver state-of-the-art system efficiency.
What carries the argument
Global ranking of output-feature importance across the full model to decide per-feature bit widths, together with two-step dequantization that maps to tensor cores and a pipeline that hides conversion latency.
Load-bearing premise
Ranking output features by their global importance across layers produces a quantization configuration whose accuracy-memory trade-off is meaningfully better than existing per-layer heuristics, and the ranking itself can be computed without prohibitive extra cost or data.
What would settle it
Reproduce the reported perplexity and MMLU-Pro numbers for Llama 3.1 70B at the same average bit width as the prior SOTA mixed-precision baseline and check whether the gap shrinks from roughly 0.5 to 0.2 in perplexity and from 1.92 to 0.99 in MMLU-Pro loss.
If this is right
- Quantized LLMs retain lower perplexity on next-token prediction tasks at the same memory budget.
- Downstream benchmark degradation on suites such as MMLU-Pro is cut by roughly half compared with prior mixed-precision methods.
- The same global allocation pattern applies across multiple model families without per-model retuning of the ranking step.
- System throughput remains competitive because the dequantization and pipeline hide the overhead of the mixed data types.
Where Pith is reading between the lines
- The global ranking step could be computed once on a calibration set and reused for fine-tuning or continued pre-training of the same model family.
- If the importance ranking correlates with activation magnitude or gradient sensitivity, similar global ordering might improve other compression methods such as pruning or low-rank adaptation.
- Hardware vendors might prioritize native support for the two-step dequantization pattern if it becomes common in deployed quantized models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MixLLM, a mixed-precision quantization method for LLMs that assigns bit-widths to output features according to their global importance across the model (rather than per-layer statistics). It includes algorithm-system co-design with two-step dequantization, fast data-type conversion, and a software pipeline to overlap memory access, dequantization, and MatMul. Experiments claim that adding only 10% more bits reduces perplexity degradation from ~0.5 to within 0.2 on Llama 3.1 70B and MMLU-Pro loss from 1.92 to 0.99 versus three SOTA baselines, while also achieving state-of-the-art system efficiency. Code is released.
Significance. If the global-ranking approach demonstrably outperforms per-layer heuristics without prohibitive overhead or scale bias, the work would strengthen mixed-precision quantization for large models by better allocating limited bits. The system optimizations for Tensor Core utilization and pipelining would add practical deployment value. Releasing code supports reproducibility and is a clear strength.
major comments (2)
- [Abstract] Abstract: the central accuracy claims rest on a global importance ranking of output features, yet the text provides no description of the importance metric (e.g., activation norm, sensitivity), whether cross-layer normalization is applied, and no ablation isolating the global-versus-per-layer choice; without these the superiority over existing mixed-precision heuristics cannot be assessed.
- [Abstract] Abstract: the reported perplexity and MMLU-Pro improvements are given without error bars, without naming the three SOTA baselines or the three popular models, and without details on how the 10% bit overhead is measured, rendering the quantitative comparison unevaluable from the provided text.
minor comments (1)
- [Abstract] The abstract refers to 'three popular models' and 'three SOTA baselines' without explicit names or citations.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We agree the abstract is overly concise and will revise it to include the requested details on the importance metric, ablations, baselines, models, error bars, and bit-overhead measurement. Point-by-point responses are below.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central accuracy claims rest on a global importance ranking of output features, yet the text provides no description of the importance metric (e.g., activation norm, sensitivity), whether cross-layer normalization is applied, and no ablation isolating the global-versus-per-layer choice; without these the superiority over existing mixed-precision heuristics cannot be assessed.
Authors: We agree the abstract should briefly describe the metric. The manuscript (Section 3) defines importance via global average activation norms on a calibration set with explicit cross-layer normalization. Section 4.3 contains the requested ablation isolating global vs. per-layer ranking. We will add one sentence to the abstract summarizing the metric and citing the ablation. revision: yes
-
Referee: [Abstract] Abstract: the reported perplexity and MMLU-Pro improvements are given without error bars, without naming the three SOTA baselines or the three popular models, and without details on how the 10% bit overhead is measured, rendering the quantitative comparison unevaluable from the provided text.
Authors: We will revise the abstract to explicitly name the three SOTA baselines and the three popular models. Error bars appear in the full experimental tables and will be referenced. The 10% overhead is the relative increase in total bits versus uniform 4-bit quantization, computed from aggregate model size; we will add this definition. revision: yes
Circularity Check
No circularity; empirical method validated by direct measurements
full rationale
The paper introduces an algorithmic mixed-precision quantization scheme that ranks output features globally and assigns bit-widths accordingly, then validates the approach through perplexity and benchmark experiments on models such as Llama 3.1 70B. No derivation chain, equations, or uniqueness theorems are presented that reduce by construction to fitted parameters, self-citations, or renamed inputs. The accuracy and efficiency claims rest on reported experimental outcomes rather than any self-referential construction, satisfying the criteria for a self-contained empirical result.
Axiom & Free-Parameter Ledger
read the original abstract
Quantization has become one of the most effective methodologies to compress LLMs into smaller size. However, the existing quantization solutions still show limitations of either non-negligible accuracy drop or low system efficiency. In this paper, we propose MixLLM that explores the optimization space of mixed-precision quantization between output features, based on the insight that different features matter differently in the model. MixLLM identifies the important output features in the global view rather than within each single layer, effectively assigning larger bit-width to output features that need it the most to achieve high accuracy and low memory usage. We present the sweet spot of quantization configuration of algorithm-system co-design with high accuracy and system efficiency. To address the system challenge, we design the two-step dequantization to make use of the Tensor Core easily and fast data type conversion to reduce dequantization overhead, and present the software pipeline to overlap the memory access, dequantization and the MatMul to the best. Extensive experiments show that with only 10\% more bits, the perplexity increase can be reduced from about 0.5 in SOTA to within 0.2 for Llama 3.1 70B, while MMLU-Pro loss can be reduced from 1.92 to 0.99 over the SOTA of three popular models. Besides its superior accuracy, MixLLM also achieves state-of-the-art system efficiency. Code is released at https://github.com/microsoft/MixLLM.
Figures
Forward citations
Cited by 3 Pith papers
-
Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models
Learned per-group bit-widths yield a reusable low-bit recipe that makes language models simultaneously larger in parameters and smaller in storage than FP16 baselines, with growing decode speedups.
-
Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment
KL divergence correlates with benchmark scores over wide quantization ranges but loses all predictive power in the near-baseline silent zone because it tracks disagreement volume rather than direction.
-
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
SplitQ improves low-bit PTQ for VLMs by isolating modality-specific outlier channels via MOCD and applying dual-branch adaptive calibration via ACC, outperforming prior methods on six datasets across W4A8 to W3A2 settings.
Reference graph
Works this paper leans on
-
[1]
Hamdy Abdelkhalik, Yehia Arafa, Nandakishore Santhi, and Abdel-Hameed A. Badawy. Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis. In IEEE High Performance Extreme Computing Conference, HPEC 2022, Waltham, MA, USA, September 19- 23, 2022, pages 1–8. IEEE, 2022
work page 2022
-
[2]
Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. Taming throughput-latency tradeoff in LLM inference with sarathi-serve. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024 , pages 117–134. USENIX Ass...
work page 2024
-
[3]
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. Quarot: Outlier-free 4-bit inference in rotated llms. CoRR, abs/2404.00456, 2024
-
[4]
AutoAWQ. Autoawq. https://github.com/casper-hansen/AutoAWQ, Cited Sep. 2024
work page 2024
- [5]
-
[6]
A systematic classification of knowledge, reasoning, and context within the ARC dataset
Michael Boratko, Harshit Padigela, Divyendra Mikkilineni, Pritish Yuvraj, Rajarshi Das, Andrew Mc- Callum, Maria Chang, Achille Fokoue-Nkoutche, Pavan Kapanipathi, Nicholas Mattei, Ryan Musa, Kartik Talamadupula, and Michael Witbrock. A systematic classification of knowledge, reasoning, and context within the ARC dataset. In Eunsol Choi, Minjoon Seo, Danq...
work page 2018
-
[7]
Sparks of Artificial General Intelligence: Early experiments with GPT-4
S´ ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco T´ ulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with GPT-4. CoRR, abs/2303.12712, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[8]
Quip: 2-bit quantization of large language models with guarantees
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa. Quip: 2-bit quantization of large language models with guarantees. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, Neu...
work page 2023
-
[9]
CUTLASS. Cutlass. https://github.com/NVIDIA/cutlass, Cited Sep. 2024
work page 2024
-
[10]
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm.int8(): 8-bit matrix multipli- cation for transformers at scale. CoRR, abs/2208.07339, 2022
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[11]
Spqr: A sparse-quantized representation for near-lossless LLM weight compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. Spqr: A sparse-quantized representation for near-lossless LLM weight compression. In The Twelfth International Conference on Learning Represen- tations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . Ope...
work page 2024
-
[12]
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ: hessian aware quantization of neural networks with mixed-precision. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 , pages 293–302. IEEE, 2019
work page 2019
-
[13]
Optimal brain compression: A framework for accurate post-training quantization and pruning
Elias Frantar and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans,...
work page 2022
-
[14]
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: accurate post-training quantization for generative pre-trained transformers. CoRR, abs/2210.17323, 2022
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[15]
A framework for few-shot language model evaluation, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework...
work page 2024
- [16]
-
[17]
Qwen2.5: A party of foundation models
Alibaba Group. Qwen2.5: A party of foundation models. https://qwenlm.github.io/blog/qwen2.5/, Cited Nov. 2024
work page 2024
-
[18]
Babak Hassibi, David G. Stork, and Gregory J. Wolff. Optimal brain surgeon and general network pruning. In Proceedings of International Conference on Neural Networks (ICNN’88), San Francisco, CA, USA, March 28 - April 1, 1993 , pages 293–299. IEEE, 1993
work page 1993
-
[19]
2024.Deepspeed-fastgen: High-throughput text generation for llms via mii and deepspeed-inference
Connor Holmes, Masahiro Tanaka, Michael Wyatt, Ammar Ahmad Awan, Jeff Rasley, Samyam Ra- jbhandari, Reza Yazdani Aminabadi, Heyang Qin, Arash Bakhtiari, Lev Kurilenko, and Yuxiong He. Deepspeed-fastgen: High-throughput text generation for llms via MII and deepspeed-inference. CoRR, abs/2401.08671, 2024. 14
-
[20]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L´ elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timoth´ ee Lacroix, and William El Sayed. Mistral 7b. CoRR, abs/2310.06825, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[21]
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, and Kurt Keutzer. Squeezellm: Dense-and-sparse quantization. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024
work page 2024
-
[22]
Tanishq Kumar, Zachary Ankner, Benjamin F Spector, Blake Bordelon, Niklas Muennighoff, Mansheej Paul, Cengiz Pehlevan, Christopher R´ e, and Aditi Raghunathan. Scaling laws for precision. arXiv preprint arXiv:2411.04330, 2024
-
[23]
Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. A fast post-training pruning framework for transformers. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, ...
work page 2022
-
[24]
Yann LeCun, John S. Denker, and Sara A. Solla. Optimal brain damage. In David S. Touretzky, editor, Advances in Neural Information Processing Systems 2, [NIPS Conference, Denver, Colorado, USA, November 27-30, 1989] , pages 598–605. Morgan Kaufmann, 1989
work page 1989
-
[25]
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models. In Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan, editors, Thirty-Eighth AAAI Conference on Artificial Intelli- gence, AAAI 2024, Thirty-Sixth Conference on Innovative...
work page 2024
-
[26]
AWQ: activation-aware weight quantization for on-device LLM compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: activation-aware weight quantization for on-device LLM compression and acceleration. In Phillip B. Gibbons, Gennady Pekhimenko, and Christopher De Sa, editors, Proceedings of the Seventh Annual Conference on Machine Lea...
work page 2024
-
[27]
Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han. Qserve: W4A8KV4 quantization and system co-design for efficient LLM serving.CoRR, abs/2405.04532, 2024
-
[28]
SpinQuant: LLM quantization with learned rotations
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krish- namoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. Spinquant: LLM quantization with learned rotations. CoRR, abs/2405.16406, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[29]
Affinequant: Affine transformation quantization for large language models
Yuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling, Xuefeng Xiao, Rui Wang, Shilei Wen, Fei Chao, and Rongrong Ji. Affinequant: Affine transformation quantization for large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024
work page 2024
-
[30]
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. Shortgpt: Layers in large language models are more redundant than you expect. CoRR, abs/2403.03853, 2024
work page Pith review arXiv 2024
-
[31]
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. 15
work page 2017
-
[32]
Meta. Llama 3. https://ai.meta.com/blog/meta-llama-3, Cited Sep. 2024
work page 2024
-
[33]
MIT-Han-Lab. lmquant. https://github.com/mit-han-lab/lmquant, Cited Sep. 2024
work page 2024
-
[34]
MIT-Han-Lab. Pileval. https://huggingface.co/datasets/mit-han-lab/pile-val-backup , Cited Sep. 2024
work page 2024
-
[35]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. , 21:140:1–140:67, 2020
work page 2020
-
[36]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. CoRR, abs/2311.12022, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[37]
Omniquant: Omnidirectionally calibrated quantization for large language models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. Omniquant: Omnidirectionally calibrated quantization for large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024
work page 2024
-
[38]
Musr: Testing the limits of chain-of-thought with multistep soft reasoning
Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. Musr: Testing the limits of chain-of-thought with multistep soft reasoning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024
work page 2024
-
[39]
Mirac Suzgun, Nathan Scales, Nathanael Sch¨ arli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 13003...
work page 2023
-
[40]
TensorRT-LLM. Tensorrt-llm. https://github.com/NVIDIA/TensorRT-LLM, Cited Sep. 2024
work page 2024
-
[41]
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. CoRR, abs/2406.01574, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[42]
Xiaoxia Wu, Haojun Xia, Stephen Youn, Zhen Zheng, Shiyang Chen, Arash Bakhtiari, Michael Wy- att, Reza Yazdani Aminabadi, Yuxiong He, Olatunji Ruwase, Leon Song, and Zhewei Yao. Zero- quant(4+2): Redefining llms quantization with a new fp6-centric strategy for diverse generative tasks. CoRR, abs/2312.08583, 2023
-
[43]
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity. Proc. VLDB Endow., 17(2):211–224, 2023
work page 2023
-
[44]
Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, and Shuaiwen Leon Song. Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus. In Saurabh Bagchi and Yiying Zhang, editors, ...
work page 2024
-
[45]
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Micka¨ el Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In Andreas Krause, Emma 16 Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Interna- tional Conference on Machine Learning, ICML 2023, 23-29 Jul...
work page 2023
-
[46]
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Inf...
work page 2022
-
[47]
Orca: A distributed serving system for transformer-based generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. Orca: A distributed serving system for transformer-based generative models. In Marcos K. Aguilera and Hakim Weatherspoon, editors, 16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2022, Carlsbad, CA, USA, July 11-13, 2022 , pages 521–538. USENIX Associ...
work page 2022
-
[48]
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. RPTQ: reorder-based post-training quantization for large language models. CoRR, abs/2304.01089, 2023
-
[49]
Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Llu´ ıs M` arquez, editors,Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , pages...
work page 2019
-
[50]
Atom: Low-bit quantization for efficient and accurate LLM serving
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishna- murthy, Tianqi Chen, and Baris Kasikci. Atom: Low-bit quantization for efficient and accurate LLM serving. In Proceedings of the Seventh Annual Conference on Machine Learning and Systems, MLSys 2024, Santa Clara, CA, USA, May 13-16, 2024 , 2024
work page 2024
-
[51]
Zhen Zheng, Xin Ji, Taosong Fang, Fanghao Zhou, Chuanjie Liu, and Gang Peng. Batchllm: Optimizing large batched llm inference with global prefix sharing and throughput-oriented token batching. arXiv preprint arXiv:2412.03594, 2024. 17
work page internal anchor Pith review Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.