Pith. sign in

REVIEW 2 major objections 2 minor 1 cited by

SplitZip compresses KV caches losslessly at over 600 GB/s on GPUs using a fixed exponent codebook and sparse escape stream.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-01 00:55 UTC pith:DXFDO6JM

load-bearing objection SplitZip delivers measured throughput gains on KV transfer with a static exponent codebook, but offers no data on how well that codebook holds up across models. the 2 major comments →

arxiv 2605.01708 v3 pith:DXFDO6JM submitted 2026-05-03 cs.DC cs.AIcs.LG

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving

classification cs.DC cs.AIcs.LG
keywords KV cache compressionlossless compressiondisaggregated LLM servingGPU compressionexponent encodingprefill-decode disaggregationBF16 activation tensors
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces SplitZip to solve the transfer bottleneck when moving KV caches from prefill workers to decode workers in disaggregated LLM serving. It encodes common floating-point exponents with fixed-length codes drawn from a pre-calibrated top-16 set and routes infrequent exponents through a lightweight sparse correction stream of position-value pairs. This structure keeps both compression and decompression fast on GPUs while preserving every bit of the original tensor and requiring no changes to model execution. A reader would care because KV transfer currently limits scaling for long-context and agentic workloads, and existing lossless codecs cannot keep pace with GPU production rates. If the method works as described, serving systems can move more data per second without adding bandwidth hardware.

Core claim

SplitZip encodes KV activations by mapping frequent exponent values to fixed-length codes from a static top-16 codebook while sending rare exponents through a sparse escape stream of (position, value) pairs, producing a regular dense path and an occasional correction path that together run efficiently on GPUs and deliver exact reconstruction.

What carries the argument

SplitZip compressor combining a dense fixed-length encoding path for the top-16 exponents with a sparse escape stream for outliers.

Load-bearing premise

A single fixed top-16 exponent codebook calibrated ahead of time stays effective across different models, inputs, and workloads.

What would settle it

Measure compression throughput and end-to-end speedup when SplitZip is applied to KV tensors whose exponent histogram differs markedly from the calibration distribution; if speedups fall below the reported 1.23–1.32× range, the fixed-codebook premise does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • KV cache transfer reaches up to 1.32× speedup for BF16 tensors.
  • Time-to-first-token improves by up to 1.30× and request throughput by 1.23×.
  • The same scheme yields up to 1.14× compression on FP8 KV caches relative to native E5M2.
  • Both compression at 613.3 GB/s and decompression at 2181.8 GB/s fit inside the critical path of disaggregated serving.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The fixed-codebook design could extend to other floating-point activation tensors that share similar exponent skew.
  • Lowering KV transfer cost might relax the need for the highest-bandwidth interconnects between prefill and decode nodes.
  • Because the method requires no model changes, it could be dropped into existing serving stacks to test real-world gains without retraining.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. SplitZip introduces a GPU-friendly lossless compressor for KV caches in prefill-decode disaggregated LLM serving. It exploits redundancy in BF16 (and FP8) floating-point exponents by encoding the most frequent 16 values with fixed-length codes and routing rare exponents to a sparse escape stream of (position, value) pairs. A pre-calibrated static top-16 codebook removes the need for online histogramming. On real BF16 activation tensors the method reports 613.3 GB/s compression and 2181.8 GB/s decompression throughput, yielding up to 1.32× end-to-end KV-transfer speedup, 1.30× TTFT improvement, and 1.23× higher request throughput. The same scheme provides up to 1.14× compression over native E5M2 FP8. Public code is provided.

Significance. If the static codebook remains effective across models, layers, and workloads, the technique directly mitigates a practical bottleneck in disaggregated serving for long-context and agentic workloads. The GPU kernel design that keeps both the dense path and the sparse correction efficient, together with the open-source release, are concrete strengths that would support adoption and follow-on work.

major comments (2)
  1. [Evaluation section] Evaluation section: the central performance claims (613 GB/s compression, 1.32× transfer speedup) rest on the assumption that a single pre-calibrated top-16 exponent codebook keeps the escape stream sufficiently sparse for all evaluated tensors. No coverage statistics, per-model histogram overlap, or worst-case ratio under distribution shift are reported, leaving the robustness of the “eliminates online histogramming” advantage unquantified.
  2. [§3] §3 (method description): the paper states that the approach “integrates into existing serving frameworks without modifying model execution,” yet provides no concrete interface description or pseudocode showing how the compressor is invoked on the KV tensors produced by the prefill worker before the network transfer.
minor comments (2)
  1. [Abstract and §4] Abstract and §4: the phrase “real BF16 activation tensors” should be accompanied by the specific models, layer indices, and sequence-length ranges used to generate the reported throughput numbers.
  2. Figure captions and tables: axis labels and legend entries for the end-to-end speedup plots should explicitly state whether the baseline is uncompressed BF16 transfer or a prior lossless codec.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback on robustness quantification and integration clarity. We address both major comments below and will update the manuscript to strengthen these aspects.

read point-by-point responses
  1. Referee: [Evaluation section] Evaluation section: the central performance claims (613 GB/s compression, 1.32× transfer speedup) rest on the assumption that a single pre-calibrated top-16 exponent codebook keeps the escape stream sufficiently sparse for all evaluated tensors. No coverage statistics, per-model histogram overlap, or worst-case ratio under distribution shift are reported, leaving the robustness of the “eliminates online histogramming” advantage unquantified.

    Authors: We agree that explicit coverage statistics would strengthen the robustness claim. The static codebook was calibrated on a diverse collection of BF16 KV tensors drawn from multiple models and layers to capture common exponent distributions. Across the reported workloads the escape stream remained small enough to sustain the measured throughputs. In revision we will add a table (or subsection in Evaluation) reporting per-model top-16 coverage percentages and observed escape ratios, together with a brief note on the calibration set composition. This directly addresses the concern without requiring new experiments. revision: yes

  2. Referee: [§3] §3 (method description): the paper states that the approach “integrates into existing serving frameworks without modifying model execution,” yet provides no concrete interface description or pseudocode showing how the compressor is invoked on the KV tensors produced by the prefill worker before the network transfer.

    Authors: The design applies the compressor to the already-materialized KV tensors immediately after prefill generation and before the network send, leaving the model’s attention and KV-write logic unchanged. To make the boundary explicit we will insert a short pseudocode listing in §3 (or a new figure) showing the sequence: after the prefill worker produces the KV tensor, call compress(KV) to obtain the encoded payload for transfer; the decode worker calls decompress on receipt. This addition clarifies the integration point while preserving the “no model modification” property. revision: yes

Circularity Check

0 steps flagged

No circularity: performance claims are direct measurements, not derived predictions

full rationale

The paper reports throughput and speedup numbers as empirical measurements on real BF16 activation tensors (e.g., 613.3 GB/s compression). The top-16 codebook is a fixed, pre-calibrated design choice whose effect is measured rather than used to derive the reported figures by construction. No equations, self-citations, or fitted inputs are presented as load-bearing derivations that reduce to the inputs themselves. The work is self-contained against external benchmarks with no self-referential reduction.

Axiom & Free-Parameter Ledger

1 free parameters · 1 axioms · 0 invented entities

The approach rests on the domain assumption of exponent redundancy in KV activations and the practical choice of a static codebook. No new physical entities or unproven mathematical axioms are introduced.

free parameters (1)
  • top-16 exponent codebook
    Calibrated in advance to eliminate online histogramming; selection of the 16 values is data-dependent.
axioms (1)
  • domain assumption KV activations exhibit sufficient redundancy in floating-point exponents to make a small fixed codebook effective
    Core premise enabling the fixed-length encoding path.

pith-pipeline@v0.9.1-grok · 5861 in / 1270 out tokens · 31152 ms · 2026-07-01T00:55:15.406423+00:00 · methodology

0 comments
read the original abstract

Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to load-balance between the compute-bound prefill phase and the memory-bound decode phase. Under this design, prefill workers generate a KV cache that must be transferred to decode workers before generation can begin. With these workers residing on different physical systems, this transfer becomes a significant bottleneck to serving LLMs at scale, especially for long-input and agentic workloads. Existing lossless codecs are unsuitable here as they primarily target offline weight compression, run on CPUs, or use variable-length coding whose compression cannot keep up with KV production during prefill. We introduce SplitZip, a GPU-friendly lossless compressor for KV cache transfer that preserves KV tensors bitwise and integrates into existing serving frameworks without modifying model execution. SplitZip exploits redundancy in floating-point exponents of KV activations, encoding frequent exponent values with fixed-length codes and routing rare exponents through a sparse escape stream of (position, value). A calibrated top-16 exponent codebook eliminates online histogramming, while the regular dense path and sparse escape correction make both encoding and decoding efficient on GPUs. On real BF16 activation tensors, SplitZip achieves $613.3$ GB/s compression throughput and $2181.8$ GB/s decompression throughput, outperforming prior lossless compressors on the critical codec path. End-to-end transfer experiments show up to $1.32\times$ speedup for BF16 KV cache transfer, $1.30\times$ speedup for TTFT, and $1.23\times$ increase in Request Throughput. The same approach extends to FP8 KV caches, providing up to $1.14\times$ compression over native E5M2. Code is available at https://github.com/Intelligent-Microsystems-Lab/SplitZip

Figures

Figures reproduced from arXiv: 2605.01708 by Siddharth Joshi, Yipin Guo.

Figure 1
Figure 1. Figure 1: Overview of SplitZip for lossless KV-cache compression. SplitZip exploits the redun￾dancy of the BF16 exponent field while keeping sign and mantissa exact. It encodes the most frequent exponent values with fixed 4-bit codes and stores rare values in a small escape buffer with their positions and raw exponents. This design enables GPU-friendly parallel decoding and reconstructs the original KV cache bit-exa… view at source ↗
Figure 1
Figure 1. Figure 1: Overview of SplitZip for lossless KV cache compression. SplitZip exploits the redundancy of the BF16 exponent field while keeping sign and mantissa exact. It encodes the most frequent exponent values with fixed 4-bit codes and stores rare values in a small escape buffer with their positions and raw exponents. This design enables GPU-friendly parallel decoding and reconstructs the original KV cache bit-exac… view at source ↗
Figure 2
Figure 2. Figure 2: End-to-end speedup on Qwen3-32B across sequence-length sweeps with SplitZip enabled view at source ↗
Figure 3
Figure 3. Figure 3: KV-cache transfer time across sequence-length and batch-size sweeps using Mooncake[ view at source ↗
Figure 4
Figure 4. Figure 4: Transmission-time breakdown on Qwen3-32B. Native transfer sends the raw BF16 KV view at source ↗
Figure 5
Figure 5. Figure 5: Layer-wise coverage under a fixed shared Top-16 codebook on Qwen3-32B. < 99.0% 99.0--99.5% 99.5--99.8% 99.8% 0 16 32 48 64 Number of Layers 0 1 2 61 2 5 13 44 K cache V cache We further test whether a small number of layers ex￾hibit heavier-tailed exponent distributions that would increase the escape rate. For Qwen3-32B, we profile BF16 KV caches from all 64 layers, select one shared Top-16 codebook from t… view at source ↗
Figure 6
Figure 6. Figure 6: Escape value capture chunk-size ablation. We evaluate the sensitivity of SplitZip to the escape￾collection chunk size. The chunk size determines the local range used to record escape positions. When the chunk size is no larger than 256, each escape position can be represented with an 8-bit integer instead of a 16-bit integer, reducing the metadata cost of the sparse escape stream. However, this metadata sa… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving

    cs.LG 2026-06 unverdicted novelty 6.0

    SpectrumKV applies per-token mixed-precision KV cache transfer (FP16/INT8/INT4) with a model-specific probe for INT4 tolerance, achieving better perplexity and retrieval than PDTrim at equivalent budgets on Qwen2.5-7B...

Reference graph

Works this paper leans on

34 extracted references · 20 canonical work pages · cited by 1 Pith paper · 8 internal anchors

  1. [1]

    Quad length codes for lossless compression of e4m3, 2026

    Aditya Agrawal, Albert Magyar, Hiteshwar Eswaraiah, Patrick Sheridan, Pradeep Janedula, Ravi Krishnan Venkatesan, Krishna Nair, and Ravi Iyer. Quad length codes for lossless compression of e4m3, 2026. URL https://arxiv.org/abs/2602.17849

  2. [2]

    Single-stage huffman encoder for ml compression, 2026

    Aditya Agrawal, Albert Magyar, Hiteshwar Eswaraiah, Patrick Sheridan, Pradeep Janedula, Ravi Krishnan Venkatesan, Krishna Nair, and Ravi Iyer. Single-stage huffman encoder for ml compression, 2026. URL https://arxiv.org/abs/2601.10673

  3. [3]

    Longformer: The Long-Document Transformer

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer, 2020. URLhttps://arxiv.org/abs/2004.05150

  4. [4]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian...

  5. [5]

    Training Verifiers to Solve Math Word Problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URLhttps://arxiv.org/abs/2110.14168

  6. [6]

    Sglang.https://github.com/sgl-project/sglang, 2026

    LMSYS Corp. Sglang.https://github.com/sgl-project/sglang, 2026

  7. [7]

    Zipserv: Fast and memory-efficient llm inference with hardware-aware lossless compression

    Ruibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li, Weile Luo, Qiang Wang, Wei Wang, and Xiaowen Chu. Zipserv: Fast and memory-efficient llm inference with hardware-aware lossless compression. InProceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, V olume 2, ASPLOS ’26, page 2264–2280, Ne...

  8. [8]

    Measuring Massive Multitask Language Understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding, 2021. URL https://arxiv.org/ abs/2009.03300

  9. [9]

    Zipnn: Lossless compression for ai models, 2024

    Moshik Hershcovitch, Andrew Wood, Leshem Choshen, Guy Girmonsky, Roy Leibovitz, Ilias Ennmouri, Michal Malka, Peter Chin, Swaminathan Sundararaman, and Danny Harnik. Zipnn: Lossless compression for ai models, 2024. URLhttps://arxiv.org/abs/2411.05239

  10. [10]

    Mahoney, Sophia Shao, Kurt Keutzer, and Amir Gholami

    Coleman Richard Charles Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Sophia Shao, Kurt Keutzer, and Amir Gholami. KVQuant: Towards 10 million context length LLM inference with KV cache quantization. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URLhttps://openreview.net/forum?id=0LXotew9Du

  11. [11]

    Disaggregated inference at scale with pytorch & vllm

    Hongyi Jia, Jinghui Zhang, Lu Fang, Stephen Chen, Yan Cui, Ye (Charlotte) Qi, and Zijing Liu. Disaggregated inference at scale with pytorch & vllm. https://pytorch.org/blog/ disaggregated-inference-at-scale-with-pytorch-vllm/ , September 2025. PyTorch Blog. Ac- cessed: April 26, 2026

  12. [12]

    Flowkv: A disaggregated inference framework with low-latency kv cache transfer and load-aware scheduling.arXiv preprint arXiv:2504.03775, 2025

    Weiqing Li, Guochao Jiang, Xiangyong Ding, Zhangcheng Tao, Chuzhan Hao, Chenfeng Xu, Yuewei Zhang, and Hao Wang. Flowkv: A disaggregated inference framework with low-latency kv cache transfer and load-aware scheduling, 2025. URLhttps://arxiv.org/abs/2504.03775

  13. [13]

    SnapKV: LLM knows what you are looking for before generation

    Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. SnapKV: LLM knows what you are looking for before generation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https: //openreview.net/forum?id=poE54GOq2l

  14. [14]

    A high-throughput gpu framework for adaptive lossless compression of floating-point data, 2025

    Zheng Li, Weiyan Wang, Ruiyuan Li, Chao Chen, Xianlei Long, Linjiang Zheng, Quanqing Xu, and Chuanhui Yang. A high-throughput gpu framework for adaptive lossless compression of floating-point data, 2025. URLhttps://arxiv.org/abs/2511.04140. 10

  15. [15]

    PM-KVQ: Progressive mixed-precision KV cache quantization for long-cot LLMs

    Tengxuan Liu, Shiyao Li, Jiayi Yang, Tianchen Zhao, Feng Zhou, Xiaohui Song, Guohao Dai, Shengen Yan, Huazhong Yang, and Yu Wang. PM-KVQ: Progressive mixed-precision KV cache quantization for long-cot LLMs. InThe F ourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=Vem6FQvRvq

  16. [16]

    Pointer sentinel mixture models,

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models,

  17. [17]

    URLhttps://arxiv.org/abs/1609.07843

  18. [18]

    The Llama 3 Herd of Models

    Meta. The llama 3 herd of models, 2024. URLhttps://arxiv.org/abs/2407.21783

  19. [19]

    Dynamo.https://github.com/ai-dynamo/dynamo, 2026

    NVIDIA. Dynamo.https://github.com/ai-dynamo/dynamo, 2026

  20. [20]

    Nvcomp.https://github.com/NVIDIA/nvcomp, 2026

    NVIDIA. Nvcomp.https://github.com/NVIDIA/nvcomp, 2026

  21. [21]

    arXiv preprint arXiv:2311.18677 , year=

    Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient generative llm inference using phase splitting, 2024. URL https: //arxiv.org/abs/2311.18677

  22. [22]

    Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

    Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, and Mingxing Zhang. Prefill-as-a-service: Kvcache of next-generation models could go cross-datacenter, 2026. URL https://arxiv.org/abs/2604.15039

  23. [23]

    Mooncake: A kvcache-centric disaggregated architecture for llm serving.ACM Trans

    Qin Ruoyu, Li Zheming, He Weiran, Cui Jialei, Tang Heyi, Ren Feng, Ma Teng, Cai Shangming, Zhang Yineng, Zhang Mingxing, Wu Yongwei, Zheng Weimin, and Xu Xinran. Mooncake: A kvcache-centric disaggregated architecture for llm serving.ACM Trans. Storage, nov 2025. ISSN 1553-3077. doi: 10.1145/3773772. URLhttps://doi.org/10.1145/3773772

  24. [24]

    Kvlinc: Kv cache quantization with hadamard rotation and linear correction,

    Utkarsh Saxena and Kaushik Roy. Kvlinc: Kv cache quantization with hadamard rotation and linear correction.arXiv preprint arXiv:2510.05373, 2025

  25. [25]

    vllm.https://github.com/vllm-project/vllm, 2026

    vLLM Team. vllm.https://github.com/vllm-project/vllm, 2026

  26. [26]

    Zipllm: Efficient llm storage via model-aware synergistic data deduplication and compression, 2025

    Zirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang, and Yue Cheng. Zipllm: Efficient llm storage via model-aware synergistic data deduplication and compression, 2025. URL https://arxiv.org/abs/ 2505.06252

  27. [27]

    Efficient streaming language models with attention sinks

    Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=NG7sS51zVF

  28. [28]

    Qwen3 Technical Report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...

  29. [29]

    Huff-llm: End- to-end lossless compression for efficient llm inference,

    Patrick Yubeaton, Tareq Mahmoud, Shehab Naga, Pooria Taheri, Tianhua Xia, Arun George, Yasmein Khalil, Sai Qian Zhang, Siddharth Joshi, Chinmay Hegde, and Siddharth Garg. Huff-llm: End-to-end lossless compression for efficient llm inference, 2025. URLhttps://arxiv.org/abs/2502.00922

  30. [30]

    KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization

    Tianyi Zhang, Jonah Wonkyu Yi, Zhaozhuo Xu, and Anshumali Shrivastava. KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum? id=pNnvzQsS4P

  31. [31]

    70% size, 100% accuracy: Lossless LLM compression for efficient GPU inference via dynamic-length float (DFloat11)

    Tianyi Zhang, Mohsen Hariri, Shaochen Zhong, Vipin Chaudhary, Yang Sui, Xia Hu, and Anshumali Shrivastava. 70% size, 100% accuracy: Lossless LLM compression for efficient GPU inference via dynamic-length float (DFloat11). InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URLhttps://openreview.net/forum?id=xdNAVP7TGy

  32. [32]

    H2o: Heavy-hitter oracle for efficient generative inference of large language models

    Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Re, Clark Barrett, Zhangyang Wang, and Beidi Chen. H2o: Heavy-hitter oracle for efficient generative inference of large language models. InThirty-seventh Conference on Neural Information Processing Systems, 2023. URLhttps://openreview.net/...

  33. [33]

    Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving

    Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. Distserve: disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Conference on Operating Systems Design and Implementation, OSDI’24, USA, 2024. USENIX Association. ISBN 978-1-939133-40-3

  34. [34]

    Head-driven phrase structure grammar parsing on penn treebank, 2020

    Junru Zhou and Hai Zhao. Head-driven phrase structure grammar parsing on penn treebank, 2020. URL https://arxiv.org/abs/1907.02684. 12 A Quad64 V ectorized Encoding Table 8: Quad64 vectorization ablation. Variant Encode Speedup GB/s Scalar pair encode 416.21.00× Quad64 encode 613.31.47× We further ablate the low-level GPU encoding layout. The baseline los...