Pith. sign in

REVIEW 2 major objections 1 minor 3 cited by

MixLLM assigns bit widths to LLM output features by global importance rather than per layer to cut accuracy loss at modest extra memory cost.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-05-23 06:53 UTC

load-bearing objection MixLLM's global output-feature ranking for mixed-precision quantization is the claimed novelty but the abstract supplies no method details, ablations, or normalization steps, so the accuracy gains cannot be assessed yet. the 2 major comments →

arxiv 2412.14590 v2 submitted 2024-12-19 cs.LG

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

classification cs.LG
keywords LLM quantizationmixed-precisionoutput featuresglobal importance rankingdequantization pipelinetensor core utilizationperplexityMMLU-Pro
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper claims that different output features contribute unequally to model performance, so a single global ranking of their importance across all layers allows better allocation of higher bit widths to the features that need them most. This mixed-precision scheme is paired with hardware-aware dequantization steps and a pipeline that overlaps memory movement, conversion, and matrix multiplication on tensor cores. The result is reported as lower degradation on language modeling and downstream tasks while maintaining competitive throughput. A reader would care because existing mixed-precision methods either incur larger accuracy drops or fail to map efficiently to current accelerators.

Core claim

MixLLM identifies the important output features in the global view rather than within each single layer, effectively assigning larger bit-width to output features that need it the most to achieve high accuracy and low memory usage, while the two-step dequantization and overlapping software pipeline deliver state-of-the-art system efficiency.

What carries the argument

Global ranking of output-feature importance across the full model to decide per-feature bit widths, together with two-step dequantization that maps to tensor cores and a pipeline that hides conversion latency.

Load-bearing premise

Ranking output features by their global importance across layers produces a quantization configuration whose accuracy-memory trade-off is meaningfully better than existing per-layer heuristics, and the ranking itself can be computed without prohibitive extra cost or data.

What would settle it

Reproduce the reported perplexity and MMLU-Pro numbers for Llama 3.1 70B at the same average bit width as the prior SOTA mixed-precision baseline and check whether the gap shrinks from roughly 0.5 to 0.2 in perplexity and from 1.92 to 0.99 in MMLU-Pro loss.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Quantized LLMs retain lower perplexity on next-token prediction tasks at the same memory budget.
  • Downstream benchmark degradation on suites such as MMLU-Pro is cut by roughly half compared with prior mixed-precision methods.
  • The same global allocation pattern applies across multiple model families without per-model retuning of the ranking step.
  • System throughput remains competitive because the dequantization and pipeline hide the overhead of the mixed data types.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The global ranking step could be computed once on a calibration set and reused for fine-tuning or continued pre-training of the same model family.
  • If the importance ranking correlates with activation magnitude or gradient sensitivity, similar global ordering might improve other compression methods such as pruning or low-rank adaptation.
  • Hardware vendors might prioritize native support for the two-step dequantization pattern if it becomes common in deployed quantized models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper proposes MixLLM, a mixed-precision quantization method for LLMs that assigns bit-widths to output features according to their global importance across the model (rather than per-layer statistics). It includes algorithm-system co-design with two-step dequantization, fast data-type conversion, and a software pipeline to overlap memory access, dequantization, and MatMul. Experiments claim that adding only 10% more bits reduces perplexity degradation from ~0.5 to within 0.2 on Llama 3.1 70B and MMLU-Pro loss from 1.92 to 0.99 versus three SOTA baselines, while also achieving state-of-the-art system efficiency. Code is released.

Significance. If the global-ranking approach demonstrably outperforms per-layer heuristics without prohibitive overhead or scale bias, the work would strengthen mixed-precision quantization for large models by better allocating limited bits. The system optimizations for Tensor Core utilization and pipelining would add practical deployment value. Releasing code supports reproducibility and is a clear strength.

major comments (2)
  1. [Abstract] Abstract: the central accuracy claims rest on a global importance ranking of output features, yet the text provides no description of the importance metric (e.g., activation norm, sensitivity), whether cross-layer normalization is applied, and no ablation isolating the global-versus-per-layer choice; without these the superiority over existing mixed-precision heuristics cannot be assessed.
  2. [Abstract] Abstract: the reported perplexity and MMLU-Pro improvements are given without error bars, without naming the three SOTA baselines or the three popular models, and without details on how the 10% bit overhead is measured, rendering the quantitative comparison unevaluable from the provided text.
minor comments (1)
  1. [Abstract] The abstract refers to 'three popular models' and 'three SOTA baselines' without explicit names or citations.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We agree the abstract is overly concise and will revise it to include the requested details on the importance metric, ablations, baselines, models, error bars, and bit-overhead measurement. Point-by-point responses are below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central accuracy claims rest on a global importance ranking of output features, yet the text provides no description of the importance metric (e.g., activation norm, sensitivity), whether cross-layer normalization is applied, and no ablation isolating the global-versus-per-layer choice; without these the superiority over existing mixed-precision heuristics cannot be assessed.

    Authors: We agree the abstract should briefly describe the metric. The manuscript (Section 3) defines importance via global average activation norms on a calibration set with explicit cross-layer normalization. Section 4.3 contains the requested ablation isolating global vs. per-layer ranking. We will add one sentence to the abstract summarizing the metric and citing the ablation. revision: yes

  2. Referee: [Abstract] Abstract: the reported perplexity and MMLU-Pro improvements are given without error bars, without naming the three SOTA baselines or the three popular models, and without details on how the 10% bit overhead is measured, rendering the quantitative comparison unevaluable from the provided text.

    Authors: We will revise the abstract to explicitly name the three SOTA baselines and the three popular models. Error bars appear in the full experimental tables and will be referenced. The 10% overhead is the relative increase in total bits versus uniform 4-bit quantization, computed from aggregate model size; we will add this definition. revision: yes

Circularity Check

0 steps flagged

No circularity; empirical method validated by direct measurements

full rationale

The paper introduces an algorithmic mixed-precision quantization scheme that ranks output features globally and assigns bit-widths accordingly, then validates the approach through perplexity and benchmark experiments on models such as Llama 3.1 70B. No derivation chain, equations, or uniqueness theorems are presented that reduce by construction to fitted parameters, self-citations, or renamed inputs. The accuracy and efficiency claims rest on reported experimental outcomes rather than any self-referential construction, satisfying the criteria for a self-contained empirical result.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only review; no explicit free parameters, axioms, or invented entities are described in the provided text.

pith-pipeline@v0.9.0 · 5805 in / 1105 out tokens · 29773 ms · 2026-05-23T06:53:46.849491+00:00 · methodology

0 comments
read the original abstract

Quantization has become one of the most effective methodologies to compress LLMs into smaller size. However, the existing quantization solutions still show limitations of either non-negligible accuracy drop or low system efficiency. In this paper, we propose MixLLM that explores the optimization space of mixed-precision quantization between output features, based on the insight that different features matter differently in the model. MixLLM identifies the important output features in the global view rather than within each single layer, effectively assigning larger bit-width to output features that need it the most to achieve high accuracy and low memory usage. We present the sweet spot of quantization configuration of algorithm-system co-design with high accuracy and system efficiency. To address the system challenge, we design the two-step dequantization to make use of the Tensor Core easily and fast data type conversion to reduce dequantization overhead, and present the software pipeline to overlap the memory access, dequantization and the MatMul to the best. Extensive experiments show that with only 10\% more bits, the perplexity increase can be reduced from about 0.5 in SOTA to within 0.2 for Llama 3.1 70B, while MMLU-Pro loss can be reduced from 1.92 to 0.99 over the SOTA of three popular models. Besides its superior accuracy, MixLLM also achieves state-of-the-art system efficiency. Code is released at https://github.com/microsoft/MixLLM.

Figures

Figures reproduced from arXiv: 2412.14590 by Chuanjie Liu, Xiaonan Song, Zhen Zheng.

Figure 1
Figure 1. Figure 1: Illustration of the quantization with mixed-precision between output features and kernel execution. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The percentage of high-salient out features within each linear layer of Llama 3.1 8B model according [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The float and integer value of binary (010010110xx...x), each within a consecutive range. methodologies (e.g., GPTQ, clip search) can be applied independently to these two disjoint parts of channels. Note that we calculate the salience of the channels in one pass rather than iterative identifying the high￾salience parts in a smaller step, as we observe the single-pass method show similar results with the i… view at source ↗
Figure 4
Figure 4. Figure 4: The GPU kernel software pipeline of group-wise W4A8/W8A8 quantized MatMul. It assumes [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The speedup of two types of single linear layers over torch float16 baseline on the A100 GPU. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: The perplexity (wikitext2) of Llama 3.1 8B model with different configurations. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models

    cs.LG 2026-07 conditional novelty 7.0

    Learned per-group bit-widths yield a reusable low-bit recipe that makes language models simultaneously larger in parameters and smaller in storage than FP16 baselines, with growing decode speedups.

  2. Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment

    cs.LG 2026-06 unverdicted novelty 6.0

    KL divergence correlates with benchmark scores over wide quantization ranges but loses all predictive power in the near-baseline silent zone because it tracks disagreement volume rather than direction.

  3. Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 6.0

    SplitQ improves low-bit PTQ for VLMs by isolating modality-specific outlier channels via MOCD and applying dual-branch adaptive calibration via ACC, outperforming prior methods on six datasets across W4A8 to W3A2 settings.

Reference graph

Works this paper leans on

51 extracted references · 51 canonical work pages · cited by 3 Pith papers · 8 internal anchors

  1. [1]

    Hamdy Abdelkhalik, Yehia Arafa, Nandakishore Santhi, and Abdel-Hameed A. Badawy. Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis. In IEEE High Performance Extreme Computing Conference, HPEC 2022, Waltham, MA, USA, September 19- 23, 2022, pages 1–8. IEEE, 2022

  2. [2]

    Gulavani, Alexey Tumanov, and Ramachandran Ramjee

    Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. Taming throughput-latency tradeoff in LLM inference with sarathi-serve. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024 , pages 117–134. USENIX Ass...

  3. [3]

    Ashkboos, A

    Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. Quarot: Outlier-free 4-bit inference in rotated llms. CoRR, abs/2404.00456, 2024

  4. [4]

    AutoAWQ. Autoawq. https://github.com/casper-hansen/AutoAWQ, Cited Sep. 2024

  5. [5]

    Autogptq

    AutoGPTQ. Autogptq. https://github.com/AutoGPTQ/AutoGPTQ, Cited Sep. 2024

  6. [6]

    A systematic classification of knowledge, reasoning, and context within the ARC dataset

    Michael Boratko, Harshit Padigela, Divyendra Mikkilineni, Pritish Yuvraj, Rajarshi Das, Andrew Mc- Callum, Maria Chang, Achille Fokoue-Nkoutche, Pavan Kapanipathi, Nicholas Mattei, Ryan Musa, Kartik Talamadupula, and Michael Witbrock. A systematic classification of knowledge, reasoning, and context within the ARC dataset. In Eunsol Choi, Minjoon Seo, Danq...

  7. [7]

    Sparks of Artificial General Intelligence: Early experiments with GPT-4

    S´ ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco T´ ulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with GPT-4. CoRR, abs/2303.12712, 2023

  8. [8]

    Quip: 2-bit quantization of large language models with guarantees

    Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa. Quip: 2-bit quantization of large language models with guarantees. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, Neu...

  9. [9]

    CUTLASS. Cutlass. https://github.com/NVIDIA/cutlass, Cited Sep. 2024

  10. [10]

    LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

    Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm.int8(): 8-bit matrix multipli- cation for transformers at scale. CoRR, abs/2208.07339, 2022

  11. [11]

    Spqr: A sparse-quantized representation for near-lossless LLM weight compression

    Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. Spqr: A sparse-quantized representation for near-lossless LLM weight compression. In The Twelfth International Conference on Learning Represen- tations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . Ope...

  12. [12]

    Mahoney, and Kurt Keutzer

    Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ: hessian aware quantization of neural networks with mixed-precision. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 , pages 293–302. IEEE, 2019

  13. [13]

    Optimal brain compression: A framework for accurate post-training quantization and pruning

    Elias Frantar and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans,...

  14. [14]

    GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: accurate post-training quantization for generative pre-trained transformers. CoRR, abs/2210.17323, 2022

  15. [15]

    A framework for few-shot language model evaluation, 07 2024

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework...

  16. [16]

    Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts. The unreasonable ineffectiveness of the deeper layers. CoRR, abs/2403.17887, 2024

  17. [17]

    Qwen2.5: A party of foundation models

    Alibaba Group. Qwen2.5: A party of foundation models. https://qwenlm.github.io/blog/qwen2.5/, Cited Nov. 2024

  18. [18]

    Stork, and Gregory J

    Babak Hassibi, David G. Stork, and Gregory J. Wolff. Optimal brain surgeon and general network pruning. In Proceedings of International Conference on Neural Networks (ICNN’88), San Francisco, CA, USA, March 28 - April 1, 1993 , pages 293–299. IEEE, 1993

  19. [19]

    2024.Deepspeed-fastgen: High-throughput text generation for llms via mii and deepspeed-inference

    Connor Holmes, Masahiro Tanaka, Michael Wyatt, Ammar Ahmad Awan, Jeff Rasley, Samyam Ra- jbhandari, Reza Yazdani Aminabadi, Heyang Qin, Arash Bakhtiari, Lev Kurilenko, and Yuxiong He. Deepspeed-fastgen: High-throughput text generation for llms via MII and deepspeed-inference. CoRR, abs/2401.08671, 2024. 14

  20. [20]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L´ elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timoth´ ee Lacroix, and William El Sayed. Mistral 7b. CoRR, abs/2310.06825, 2023

  21. [21]

    Mahoney, and Kurt Keutzer

    Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, and Kurt Keutzer. Squeezellm: Dense-and-sparse quantization. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024

  22. [22]

    Scaling laws for precision,

    Tanishq Kumar, Zachary Ankner, Benjamin F Spector, Blake Bordelon, Niklas Muennighoff, Mansheej Paul, Cengiz Pehlevan, Christopher R´ e, and Aditi Raghunathan. Scaling laws for precision. arXiv preprint arXiv:2411.04330, 2024

  23. [23]

    Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami

    Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. A fast post-training pruning framework for transformers. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, ...

  24. [24]

    Denker, and Sara A

    Yann LeCun, John S. Denker, and Sara A. Solla. Optimal brain damage. In David S. Touretzky, editor, Advances in Neural Information Processing Systems 2, [NIPS Conference, Denver, Colorado, USA, November 27-30, 1989] , pages 598–605. Morgan Kaufmann, 1989

  25. [25]

    OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models

    Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models. In Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan, editors, Thirty-Eighth AAAI Conference on Artificial Intelli- gence, AAAI 2024, Thirty-Sixth Conference on Innovative...

  26. [26]

    AWQ: activation-aware weight quantization for on-device LLM compression and acceleration

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: activation-aware weight quantization for on-device LLM compression and acceleration. In Phillip B. Gibbons, Gennady Pekhimenko, and Christopher De Sa, editors, Proceedings of the Seventh Annual Conference on Machine Lea...

  27. [27]

    Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,

    Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han. Qserve: W4A8KV4 quantization and system co-design for efficient LLM serving.CoRR, abs/2405.04532, 2024

  28. [28]

    SpinQuant: LLM quantization with learned rotations

    Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krish- namoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. Spinquant: LLM quantization with learned rotations. CoRR, abs/2405.16406, 2024

  29. [29]

    Affinequant: Affine transformation quantization for large language models

    Yuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling, Xuefeng Xiao, Rui Wang, Shilei Wen, Fei Chao, and Rongrong Ji. Affinequant: Affine transformation quantization for large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024

  30. [30]

    ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

    Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. Shortgpt: Layers in large language models are more redundant than you expect. CoRR, abs/2403.03853, 2024

  31. [31]

    Pointer sentinel mixture models

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. 15

  32. [32]

    Meta. Llama 3. https://ai.meta.com/blog/meta-llama-3, Cited Sep. 2024

  33. [33]

    MIT-Han-Lab. lmquant. https://github.com/mit-han-lab/lmquant, Cited Sep. 2024

  34. [34]

    MIT-Han-Lab. Pileval. https://huggingface.co/datasets/mit-han-lab/pile-val-backup , Cited Sep. 2024

  35. [35]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. , 21:140:1–140:67, 2020

  36. [36]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. CoRR, abs/2311.12022, 2023

  37. [37]

    Omniquant: Omnidirectionally calibrated quantization for large language models

    Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. Omniquant: Omnidirectionally calibrated quantization for large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024

  38. [38]

    Musr: Testing the limits of chain-of-thought with multistep soft reasoning

    Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. Musr: Testing the limits of chain-of-thought with multistep soft reasoning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024

  39. [39]

    Le, Ed H

    Mirac Suzgun, Nathan Scales, Nathanael Sch¨ arli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 13003...

  40. [40]

    Tensorrt-llm

    TensorRT-LLM. Tensorrt-llm. https://github.com/NVIDIA/TensorRT-LLM, Cited Sep. 2024

  41. [41]

    MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. CoRR, abs/2406.01574, 2024

  42. [42]

    Zero- quant(4+2): Redefining llms quantization with a new fp6-centric strategy for diverse generative tasks

    Xiaoxia Wu, Haojun Xia, Stephen Youn, Zhen Zheng, Shiyang Chen, Arash Bakhtiari, Michael Wy- att, Reza Yazdani Aminabadi, Yuxiong He, Olatunji Ruwase, Leon Song, and Zhewei Yao. Zero- quant(4+2): Redefining llms quantization with a new fp6-centric strategy for diverse generative tasks. CoRR, abs/2312.08583, 2023

  43. [43]

    Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity

    Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity. Proc. VLDB Endow., 17(2):211–224, 2023

  44. [44]

    Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus

    Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, and Shuaiwen Leon Song. Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus. In Saurabh Bagchi and Yiying Zhang, editors, ...

  45. [45]

    Smoothquant: Accurate and efficient post-training quantization for large language models

    Guangxuan Xiao, Ji Lin, Micka¨ el Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In Andreas Krause, Emma 16 Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Interna- tional Conference on Machine Learning, ICML 2023, 23-29 Jul...

  46. [46]

    Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

    Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Inf...

  47. [47]

    Orca: A distributed serving system for transformer-based generative models

    Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. Orca: A distributed serving system for transformer-based generative models. In Marcos K. Aguilera and Hakim Weatherspoon, editors, 16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2022, Carlsbad, CA, USA, July 11-13, 2022 , pages 521–538. USENIX Associ...

  48. [48]

    Rptq: Reorder-based post-training quantization for large language models.arXiv preprint arXiv:2304.01089, 2023

    Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. RPTQ: reorder-based post-training quantization for large language models. CoRR, abs/2304.01089, 2023

  49. [49]

    Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Llu´ ıs M` arquez, editors,Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , pages...

  50. [50]

    Atom: Low-bit quantization for efficient and accurate LLM serving

    Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishna- murthy, Tianqi Chen, and Baris Kasikci. Atom: Low-bit quantization for efficient and accurate LLM serving. In Proceedings of the Seventh Annual Conference on Machine Learning and Systems, MLSys 2024, Santa Clara, CA, USA, May 13-16, 2024 , 2024

  51. [51]

    BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

    Zhen Zheng, Xin Ji, Taosong Fang, Fanghao Zhou, Chuanjie Liu, and Gang Peng. Batchllm: Optimizing large batched llm inference with global prefix sharing and throughput-oriented token batching. arXiv preprint arXiv:2412.03594, 2024. 17