Pith. sign in

Paper Citation Record · LEDGER

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

As of 23 July 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.01733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01733 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T15:17:52.966144Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact28
  • verified fuzzy17
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f57687bf-c0dd-4370-85cd-f6b593a5903f · outbound

This paper cites GPT-4 Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.660457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:36adc89dcc8f1e0f4f6db5cd3a90b1f72001236361d1bd1a6105e126d163ed88

Observation c14b926b-2300-44fb-a8c9-dc6991f28493 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.660620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:2d6d8ba050c1b396e9eb609dacf0d4cfe7459f2fec8bba75e5569c826fdd5c3f

Observation e0a4d62e-4afe-47d2-8805-8271eb57c0d5 · outbound

This paper cites DeepSeek-V3 Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving DeepSeek-V3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.658207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:f7c5705481af8d3b7eaa508def7d3177dec98d00007d7e4cf61b9d465c2b26a6

Observation eebeee84-2976-4e6f-a23b-273115040c83 · outbound

This paper cites Gemma 3 Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Gemma 3 Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.627001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:93263d46e1606e0ffc8d616f454583113976ad69a5f7889c43865b28defe06d1

Observation 6b030ef0-c96b-4223-b2b0-11a820079f2e · outbound

This paper cites Qwen3 Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Qwen3 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.650485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:b3396684f14c8d4a2c565341b4f581930c382ecbc6a3f2db28934423b875b101

Observation e31fa4ac-17ff-4e38-9ed5-fc0c8c32a2c2 · outbound

This paper cites On decoder-only architecture for speech-to-text and large language model integration,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving On decoder-only architecture for speech-to-text and large language model integration,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.748192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:968b94e0d5c505280e0b5d46a397c5e7c0618f2f136ca72bfe1d3aff012a8a7d

Observation c6928632-7829-4ac6-914f-87605b590ba5 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Moshi: a speech-text foundation model for real-time dialogue

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.651830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:6c871d68eca24273bde66911c9ca1afd4ce2d22cc65034d33689fe652bd3eddc

Observation c29d5d96-2aab-4c0c-ab5e-b27a92a4c594 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving SALMONN: Towards generic hearing abilities for large language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.746372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:7bce2b687e2cc8fa8d896575b3f01211452e6edd7306d449d0b72734050e1eb2

Observation f9023051-45ca-4da1-873d-621a9a42150e · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.630613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:eefacedea7dfd9b7a8eecff410ce3f24b7a00e47719b20c873031343cf9d0947

Observation be537fd5-64fa-4663-ab30-6a2d14204140 · outbound

This paper cites Kimi-Audio Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Kimi-Audio Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.663097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:468aae34655f0967fd222a349ec3436878a0aef75cdb1fab4bda266d09710df4

Observation 0eabac8e-4d8b-4533-9cd6-a8a2bd50c0c7 · outbound

This paper cites Step-Audio 2 Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Step-Audio 2 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.611833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:21276ab778161dc1c94f73e2d0c9e92a5c77fd4e9d20c96be10263135ff10b7a

Observation 0607b9e9-9c57-4a93-8989-ad7adb3987e7 · outbound

This paper cites Qwen3-Omni Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Qwen3-Omni Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.645879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:4893674533ffe39231430d60345a375f173de6dc522b7d9b6710e1519a5e86fd

Observation 37054825-977e-43ca-9461-dfec0a768d26 · outbound

This paper cites SLM-S2ST: A multimodal language model for direct speech- to-speech translation,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving SLM-S2ST: A multimodal language model for direct speech- to-speech translation,

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.640093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:334f3ed8df6c559c337a6d0041ed34e8595bf0f18f2e5cc6e337b30579b8141b

Observation dc79e153-5fc8-4d81-a68f-7c8ee5c5e2e1 · outbound

This paper cites Fun-audio-chat technical report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Fun-audio-chat technical report

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.618087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:2cf51d4620759c250a82292641408a8738a494f514013674a96d48a1d04e5a86

Observation acc687e0-3e9e-45ef-bc4b-ed5ef2f99f97 · outbound

This paper cites FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.665746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:32194f35167544ed98d22e5939260010cc5561d77feee08395f0e7372da65b89

Observation e18ee8f2-5b9b-449a-9983-987079b5e7ee · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.637916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:06bb8c386ec2bc940a6eea775bc4894c121df2c589de5b94c9f4829c5246fb52

Observation baac7d1e-82d0-4cc3-8c7a-5cbd29965572 · outbound

This paper cites Fun-ASR technical report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Fun-ASR technical report

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.614641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:10dd54a92b7b8a99b9aaa615cf5be1d78eb7a4a42c6318df5ce08905f5ce619f

Observation 97134e9f-f059-4a2d-b161-c1ef3e4f075f · outbound

This paper cites Index-asr technical report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Index-asr technical report

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.611719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:ee06c8c6a39864ec7f1395bb0485adfcdbcff80b1d6dcd5598e103d0b3a252ac

Observation b43b9f70-3b2c-4534-9adb-97d9ca5a04ec · outbound

This paper cites Speech recognition meets large language model: Benchmarking, models, and exploration,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Speech recognition meets large language model: Benchmarking, models, and exploration,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.763041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:9331f29e20a84ac60be78b6fc97e50d1860892747afb05dabd13debfea210591

Observation 2fe4713a-4cc0-4c1f-b756-3dd2136996bf · outbound

This paper cites Efficient Scaling for LLM-based ASR.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Efficient Scaling for LLM-based ASR

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.666034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:523a3a564ef36cd7df95ec3ea9059fa384915cf122670609350c66bf6f0925a2

Observation dc0ac3f7-8d2d-4672-aea4-6602281cda39 · outbound

This paper cites Qwen3-ASR Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Qwen3-ASR Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.674089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:a89456920b81c04f955b8e5cb8eb6c3aea54b8cb06b846488750bd465e91f62b

Observation ebfd6284-540c-4606-99ae-a97fd732d2ee · outbound

This paper cites Transducer-Llama: Integrating LLMs into streamable transducer-based speech recognition,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Transducer-Llama: Integrating LLMs into streamable transducer-based speech recognition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.751081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:74c1f644211a8e4d4ea8fcff229afba777dbb90704fa77de98e1a1d09a78b064

Observation 98c3d0c3-1d63-4b9d-a12f-d15df855b221 · outbound

This paper cites Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.627821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:956566ae4bc071599875b28e320ff348bac421402f84191c10bb06594685571d

Observation 7d9bec23-5f67-4efb-aea1-38cc4007313e · outbound

This paper cites Train short, infer long: Speech-llm enables zero-shot streamable joint asr and di- arization on long audio.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Train short, infer long: Speech-llm enables zero-shot streamable joint asr and di- arization on long audio

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.620984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:f42550a2bb09bd9d3c4424e066ee487cf3b061edfd3572848e855afd72536e41

Observation 6873758e-e06e-483d-99fa-95bc7cb828e4 · outbound

This paper cites Rlbr: Reinforcement learning with biasing rewards for contextual speech large language models,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Rlbr: Reinforcement learning with biasing rewards for contextual speech large language models,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.668833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:1819f76f9b5bfc4f0ef4e59d9a96d98d0620b8c043b91276f6d287c6fc0d7eda

Observation 0094fa6d-1b31-4dbc-a700-7eaf59c40f09 · outbound

This paper cites Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.773983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:7b6ccddf57c17beabceef29d28e0585bbb4d9801e08ebd6d61f1301449ab8853

Observation 77991ea3-a50e-4d17-bde4-23c4292930af · outbound

This paper cites Alignformer: Modality matching can achieve better zero-shot instruction-following speech-LLM,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Alignformer: Modality matching can achieve better zero-shot instruction-following speech-LLM,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.778322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:415b0fdaf0e0ad6932d187ef6052759803e6e1188ed8f2c533bc39159b4d8352

Observation 9bc29d3b-2f9f-481c-bfbf-79b0afc3686f · outbound

This paper cites Qwen2.5-Omni Technical Report.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Qwen2.5-Omni Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.637099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:fcdc77c20adbfb5fc7630f67af8d767501ec81bf98c8c3cf61f8e456cf61a4e3

Observation fe7dcb65-5b71-4ae1-8988-2dca377b95f5 · outbound

This paper cites Voxtral.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Voxtral

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.663262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:bb1bf8188724c61d492f472d9842e532d508100dd96c1c27d95955e2f5ffe8cc

Observation dc8365fa-c3df-49c8-88a1-862a96b22064 · outbound

This paper cites SpiRit- LM: Interleaved spoken and written language model,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving SpiRit- LM: Interleaved spoken and written language model,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.780437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:86b4bcf4fcefe24038072e126b474d55d12cce175137203819f2114ba7b3fc34

Observation e1fd989b-9ad9-4711-90a9-02e36a14790e · outbound

This paper cites Available: https://aclanthology.org/2025.tacl-1.2/.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Available: https://aclanthology.org/2025.tacl-1.2/

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.767355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:6144704940b96ebd39da7bbdd6f6e9d26155771563e4e8e609b185430bdf7edd

Observation ed1c8b03-cc89-4e2c-9120-56f69f53ecad · outbound

This paper cites Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.648961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:bc8e3d6050e163d4fa858c585643550079b510bbf423e0d91eda384376a3f880

Observation 8b3f074c-f435-42a4-9c1f-2dae65772be2 · outbound

This paper cites Injecting text in self-supervised speech pretraining,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Injecting text in self-supervised speech pretraining,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.765173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:3505cd54685df6a48f265318590f7e7184798300ab3ee26950ed21e275e37c36

Observation 21415aeb-b8a2-45dd-b510-86247b70f52e · outbound

This paper cites SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.624845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:1c2219e5e1cb0572e72cbc759857af94f7595704ad9cf91e5833d6a810745725

Observation 8a97ec94-c0d9-4817-b836-b62020c01b01 · outbound

This paper cites SpeechT5: Unified-modal encoder- decoder pre-training for spoken language processing,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving SpeechT5: Unified-modal encoder- decoder pre-training for spoken language processing,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.763220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:caab47ed989f7a76948d18932686e330f5d63895ad47e75e9965d5111f12b332

Observation 3a4a95b5-2717-4677-8651-6b50b3fe1c73 · outbound

This paper cites SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.621184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:41d71ce4272a16da05fa0d5db5ad608b7c0441adfc7c1a559d4460bfcc0675d0

Observation cfda2c70-b20e-4817-a848-9eb39bbaa4f0 · outbound

This paper cites JOIST: A joint speech and text streaming model for ASR,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving JOIST: A joint speech and text streaming model for ASR,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.761053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:f47936a35d62ef6d4f7d710352f5bcd22335139352aab90c089a63b511ee9280

Observation a186f019-b909-431d-9b18-7cd3a0f7e4bd · outbound

This paper cites Joint unsupervised and supervised training for multilingual ASR,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Joint unsupervised and supervised training for multilingual ASR,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.775930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:cdd12c3c56c026392b5f56fe6071ff4ea69917645dc7f96859c2345cf09d50d7

Observation c86a68ee-96a5-42c8-88e1-b5c91b4d4447 · outbound

This paper cites FastInject: Injecting unpaired text data into CTC-based ASR training,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving FastInject: Injecting unpaired text data into CTC-based ASR training,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.769482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:6d8a33efd9b77355e8e087d05d864180d1b3aea1ba46a99dc159e577aff0ede6

Observation c141dfe8-06de-4a95-80b0-92621b7dc88b · outbound

This paper cites Multitask training with text data for end-to-end speech recognition,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Multitask training with text data for end-to-end speech recognition,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.771933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:4ce993566c080f4c945707b98dc0b0be89dde1d7f1f6fd5251743914ba1571a5

Observation 73544a1e-b578-4357-9f5b-ad08ee41205c · outbound

This paper cites An attention-based joint acoustic and text on-device end-to-end model,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving An attention-based joint acoustic and text on-device end-to-end model,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.758901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:76bbdcd9982273b21939e3f054c61bb43b47bdd471d936dcf9682d770bac5368

Observation 4c4564b0-70a4-49b4-b86a-137e4b02e2ae · outbound

This paper cites Maestro: Matched speech text representations through modality matching,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Maestro: Matched speech text representations through modality matching,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.760850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:71d29261b7b406528f91b8be94dec2f5602f9955f92ad0ca75b73fa18abf6a40

Observation da805f65-5027-46ca-9a29-cb9a8717769f · outbound

This paper cites Improving joint speech-text repre- sentations without alignment,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Improving joint speech-text repre- sentations without alignment,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T06:30:44.782639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:c77043fbdcb96c1f48c808f15ca8caada964f7340bc55b80faf5805389218e63

Observation 32ec8a36-d24d-4a32-b485-a94b63aa6d2a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Measuring Massive Multitask Language Understanding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:18:32.671551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:99295a66ff8a5b8c542f54e0c14676b9e5fd72fb667ebd26354b055d99f81913

Observation 4cd0231e-454d-43fa-a999-c319ff7b00a3 · outbound

This paper cites Ultraeval-audio: A unified framework for comprehensive evaluation of audio foundation models,.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Ultraeval-audio: A unified framework for comprehensive evaluation of audio foundation models,

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.654764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:1c5a24d3ba6a9f1014fec2d896661ac8e1853ed2f0e02740003ea009157dac26

Pith citing papers

No inbound Pith citation observations are available.