Pith. sign in

Paper Citation Record · LEDGER

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

As of 22 July 2026, this Paper Citation Record lists 74 of 74 outbound references and 3 inbound Pith citation observations for arXiv:2605.00416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00416 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:05:47.128354Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:59:42.033560Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact32
  • verified fuzzy36
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c462da4b-e6a1-491f-822f-04f119b74306 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.301652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:3535322cec5efe7245de923c8f6546980304c982a85404b43984da912536b278

Observation fabb3f4d-e4b6-409e-a8bd-5b33f29847c7 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.019797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:85cd0dfbdeee049caffcf27cd1440e44f4aac18039f9c6486fae22e993c682c7

Observation 8d8c5cda-965c-4867-b44b-61ee5255f588 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.328611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e1827d35239ba76c8fa216e29eeb7ff7fbe115c568dbd2b7aea6c698b1c46946

Observation 0f33576f-9359-43f5-9b31-6b0bc32819a6 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.244008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:25da0f48f5846ced8518ba0b4cf58a4574bf92b74b01d98dbe31fbdabc963c39

Observation b8ccb53a-a27c-48d8-ab89-17ed539b5e6b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.376300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:47d14fc571f0d172af72968d2d323e5420e6182c8e29cdec4fdd6865be399391

Observation 0cf90017-53d0-4d79-9ef2-0467b0c9982f · outbound

This paper cites π 0.5: A vision-language-action model with open-world gener- alization.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies π 0.5: A vision-language-action model with open-world gener- alization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.967936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d7da05193dba7978993045723cb4c38092652747ab07bb90d21e78de8155dafd

Observation 7d87ea34-b387-4550-b3d8-7ee1509c7d5b · outbound

This paper cites Hg-dagger: Interactive imitation learning with human experts.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hg-dagger: Interactive imitation learning with human experts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.004270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e79e604f28921bc8b8bec8ce7d3d3b1305620778d996141a39050a233d46d579

Observation 17895e33-1613-4c06-8426-9117b50be000 · outbound

This paper cites Q-learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Q-learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.010141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2b2f2b84c4d009b132c121642b13f7b954754ff8266b5082980c43656aec78b0

Observation da354dc1-2e82-43ce-a670-332ebb0a43bd · outbound

This paper cites Addressing func- tion approximation error in actor-critic methods.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Addressing func- tion approximation error in actor-critic methods

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.012138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c1d62c1a58f8880e6b373d6d9f726e866450f3d6051e243f5ee5bb1ac61e13b8

Observation 4de9d740-0529-4809-9502-363d822f9f6c · outbound

This paper cites Contin- uous control with deep reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Contin- uous control with deep reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.974588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ddd8d7b0ba2942a89347fba277084d7d5184b871913b4114fae3a5c1121a58c1

Observation 20e703b6-0584-4855-b5fe-6e5dd6f1b146 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.965855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:5ffca1ad092b572a094cc334419049aa2847d202183080d57c0243b9ac262dc1

Observation 8655e183-df38-424e-a622-1e62cb0f2392 · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rl-100: Performant robotic manipulation with real-world reinforcement learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.288694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f91dde489ebcec124f86a7ac2fc7dc91348a628c26db5d0fa086dcbae899a623

Observation 789e75af-6c26-4ef8-807f-f90ae2a241d3 · outbound

This paper cites Gr-rl: Going dexterous and precise for long-horizon robotic manipulation.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Gr-rl: Going dexterous and precise for long-horizon robotic manipulation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.372068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:abe67422b8ff72d432583d95f3852839d00429d7b31d7c4a611573240a7e9014

Observation c4d77d0f-fb4d-4a20-b8e7-6f69cf9e9c29 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.386306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c168b1d45ebaf50217e99c263333b633a050c44910ea9ba23d05c86a04576053

Observation 31042599-b993-42b1-9b4f-31a07958c008 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:16b1afd52aa74e9707b7f9a427732b31732a2ea461cf6d54b804ed3caaa60982

Observation 47a58aa5-fdc0-46bd-a8af-21bfd2e46dd3 · outbound

This paper cites Serl: A software suite for sample-efficient robotic reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Serl: A software suite for sample-efficient robotic reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.997970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f9cbc52b7b5ce8176020cab30192dca254c5d47a08c82b4bbdbbd10241e15674

Observation b99c4d50-6305-4146-b423-8bda435368b0 · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.944097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:dd18bd21e318ae8a4359dc887c076b4d315056f4a7feeea539b599d7bd652e64

Observation f4b8ac92-607d-4c3c-97a1-5471416adf4d · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.305693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2ed1be6dcb8a31d6bb1b82616e2d82e993d9853c30f09db7c020b9f53832e676

Observation 4d23dbfa-1a95-47b7-8b53-3123caaf6da4 · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Interactive Post-Training for Vision-Language-Action Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.281810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:12bd57db50a2aee36a053c3a81773ac4570b40a840d7419b0ceaca1ce549d418

Observation 358de7c6-da13-4340-a79c-dbdfb6df5f8f · outbound

This paper cites pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.254193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c70d71736c2ece80a29246b3aecd6418777a8061a280d11ca3571c74ec96fba1

Observation 482c806c-1c8e-4c21-b136-39aab369ef48 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow-GRPO: Training Flow Matching Models via Online RL

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.314277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:6ce4810144f629c39a6175c03f9024245e3a9bac2a5a4963a30e1fa62562a3d1

Observation 3ad599a0-9b5c-4174-99bc-f1362d657a22 · outbound

This paper cites arXiv preprint arXiv:2505.22094 , year=.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies arXiv preprint arXiv:2505.22094 , year=

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:32.234520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:4d600335d96fb8e3bf2d0c4dc0d86d0e423fb91bb3e0f62302d57543e7dc508d

Observation d50a8dca-f1c4-4c76-8f91-55a8d8b8ae43 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Reinforcement Learning with Implicit Q-Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.277366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:cd67c79e08a310da4bd5be685655e4b62f6797f060ecf558f6a590fa31457507

Observation 109a64ea-7998-4f52-8b58-88895af6c5ef · outbound

This paper cites Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.357386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:3d45c85625d4cfd77a67db4d174cfda173e1697e6cc978633873f20adc6b40f2

Observation dad920bd-2f6b-4285-afa8-49afeca480a0 · outbound

This paper cites Q-learning with Adjoint Matching.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Q-learning with Adjoint Matching

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.273179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:3233dc49c9eac5803dcd8ed481208e1590e5726d08c5b20515eefa80243d39b8

Observation 22a6d22b-569d-45e3-bd86-0be02f5443cb · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.014084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2ff2ba344286f103963dc3af341875c0e5d9622c0427f20dec860fa49e19dc20

Observation 8446b9d6-3128-4dda-b251-182de97f9328 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.297095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:185f5d7aa37f5e6e3ff512e5b8deef96c4c8e6ce8461304496402d6f4363e412

Observation 183d3be4-3e32-49e0-9f9b-a82ef44e7052 · outbound

This paper cites Rlinf-vla: A unified and efficient framework for vla+ rl training.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rlinf-vla: A unified and efficient framework for vla+ rl training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.332868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:560f3a2a44b56aabb14e4dae3e098925b44ae4a42f258b9566da1c2016db16bf

Observation c2c2c579-7846-4540-b5f3-5da937635857 · outbound

This paper cites What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.353236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:a5c38dd908c229ba5921c0aa0e22a1097765c368054526b21b3e44aecfb3de50

Observation 8f56a2a4-f713-460f-b1f0-517b52dff45c · outbound

This paper cites RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:32.390805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:5f5f898864660cdd5515d9c649ebfd4cf1e320f688185632a938f51bc5219baa

Observation c668455f-d698-42e8-acbf-a0b1b81c651f · outbound

This paper cites Behavior- 1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Behavior- 1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.989557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:4944e71e55897d6335ff6783d7035813460219a568f0c0dc261d9c84cf4973c4

Observation f67119e0-be7b-4c26-8ace-554d8a076960 · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.259437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ad97525526ba9b90cdab89fa2bc3f6435462f383e868a63ad2d3c2ec13340276

Observation 35779cb1-d0c6-4e75-8dec-b4233fa0f725 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.946914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2be92f0e566460c832f1f8b355d783c9300c40340788c7472535bc9d1766e10a

Observation e9ea6f10-3405-413f-8245-9d2cdaffc4e1 · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digital twins (early version).

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Robotwin: Dual-arm robot benchmark with generative digital twins (early version)

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.993608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c3a696298d6d91cc1f7593a878db2946fb4e94b8901c6aa05d04d7c96f727721

Observation 1a7652de-82a3-4881-90b4-07dd324576d0 · outbound

This paper cites Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.239243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:62098d3c752e04e2087b28843fb60ee9bf14aea4ed9a5d467bacb48db32694ae

Observation 713d09b6-8adf-48be-9c14-3c955c9ce917 · outbound

This paper cites WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.346599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b2c53b90636a3ca41f92183de2be1b64e7e87a6736ab1d555e447ceddd9cc330

Observation d31d1aeb-cb82-4e75-a965-7ce03e0c0556 · outbound

This paper cites SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.395187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:aa3b3dd5d32d14383e94b5546b1bc550388b724d169deaa6663662b148be62a4

Observation d0872b62-2baf-42d0-a71d-be727a5aeaf6 · outbound

This paper cites Flow q-learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow q-learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.978628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:96b263db0a47a2e23a0ca8f6ebddc661c1f70415f6db9ea73b668a628671a7ee

Observation f2f8e029-05fd-47a5-9928-94a5a2510b04 · outbound

This paper cites Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.980895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:220b98bc14d484b0248f7014397be1c7b3889b84225546eab7a05be9309eb78b

Observation b5d88aa3-9402-4fad-a5f9-d67244d7aa3c · outbound

This paper cites Offline- to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline- to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.986777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:5115e86b9d1c3ca1f77468255c9ce7cd257988c03bdb44777f20368db2dc1aef

Observation 8bbff4a3-ff21-46d1-9343-e5b0b819ff65 · outbound

This paper cites Reincarnating reinforcement learn- ing: Reusing prior computation to accelerate progress.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Reincarnating reinforcement learn- ing: Reusing prior computation to accelerate progress

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.976756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:6581a9360d91485370128ac9f47b7a8d2bec6ca4e46c3063f8b48ec78a30cfd5

Observation f859429f-2929-4a77-b7bb-0e2ab517c43a · outbound

This paper cites Effi- cient online reinforcement learning with offline data.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Effi- cient online reinforcement learning with offline data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.006226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f624cdd341a1cd2ab6e82da96da1570dca1dfee253ba5d70b133528a46d51c70

Observation e8b82892-7e02-4b72-aec3-04041763d612 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.337531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:aab4a0a6c9679f792a07977d0a6dc835a15cb6b51b57708e389043bf21fa3bb3

Observation 39a0a6bc-4bbd-43fc-b36f-1e8850532b61 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.342064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:001745dbbcdf114ddb1999862e776d093194ae207ebb25a5b699de8a32675bd1

Observation d4d498a3-e9d2-4bfd-b160-d7e76b4381fa · outbound

This paper cites Steering Your Diffusion Policy with Latent Space Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.362409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:9932fd2dac9af6866079e965f1430d1dddbada8bccea74061904540966317b90

Observation 66da4269-1d56-4946-8c15-7a1729571d80 · outbound

This paper cites Qt-opt: Scalable deep rein- forcement learning for vision-based robotic manipula- tion.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Qt-opt: Scalable deep rein- forcement learning for vision-based robotic manipula- tion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.963482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c311e3362b73d181f0e973374703c1ba52716669074c3553757f6baf774643e4

Observation 19d92e87-c654-4ea4-a22e-77542f76a7d5 · outbound

This paper cites MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.309954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f44b0a9399fc801cfef7e58beb3670e35d5f0562d41d10212d61d532f39cae1c

Observation 17c5cc86-2c1e-40aa-b398-c061d477a16f · outbound

This paper cites Pi-qt-opt: Predictive information improves multi-task robotic reinforcement learning at scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Pi-qt-opt: Predictive information improves multi-task robotic reinforcement learning at scale

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.023775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:93393137fe0602097fcd046d419221635c23bb1c68f956b331f7b59f19dafefa

Observation 902a7074-8583-4df7-9736-685928302ae9 · outbound

This paper cites Sop: A scalable online post-training system for vision-language-action models.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Sop: A scalable online post-training system for vision-language-action models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.366990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d7e9997d373678415f0f51f0ed4869a92d9af5509e14cd23f7ba69074b22bc0e

Observation 5448ea69-3db1-4919-9941-7a1bec6e497f · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.324014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:9e5821d3a7676e1a6e6591cf4c268cd44a4c6f368d751705b077f1d97e0a59ea

Observation 633d21ce-7b75-4db9-8bda-139822ac0c4b · outbound

This paper cites Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.000134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:76c842fc78e0e639389c36d3e6a3bcd33239a33d7820f254daba5298c1e9d738

Observation 37424740-b8c3-4ff7-899f-7c0f9d55b8e9 · outbound

This paper cites Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.381350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:478ecc5b1845202787cb83757e88d74323e23df37ce7a26f035ffa80aaa6bae3

Observation c63998fc-d28f-4cdd-b489-ffdc73bd3dcb · outbound

This paper cites Flow matching for generative modeling.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow matching for generative modeling

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.015969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:1aee900619565a5eda024246c6e52559c3ad6b085eb7884bd98fe0d110294de8

Observation c43499ee-d956-4d65-ab8b-c9462145ed13 · outbound

This paper cites A dis- tributional perspective on reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies A dis- tributional perspective on reinforcement learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.984857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ecc23738c34c94fbc8e92ed620d2d1da0d0148f0f3d3ebe37a5ebfb81b0d6b6a

Observation 2c291311-8889-4673-b059-44ecf5f1b6c6 · outbound

This paper cites Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.319056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:1df731e8044f125673e7ac72386e26e903d739c919333bee015a5e1ec11ce33f

Observation c74f5f11-bd6d-4522-8274-9e8f19b811fa · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.268543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2d0d6ab0495ad2d694b6777d3273d2a1b6f8663fc3bd5765f83a0f0db93b0e86

Observation d5fa6e00-a2c0-46c4-81bd-6d9de680ed09 · outbound

This paper cites Energy-weighted flow matching for offline reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Energy-weighted flow matching for offline reinforcement learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.969944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:a110482d0828949e9c80ba970b7fa46d61a820e6b51d0beb11754a3d5d36a1bc

Observation 9e25219c-802b-430f-b70e-b129003c16bb · outbound

This paper cites Gemma 3 technical report.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Gemma 3 technical report

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.971846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:13446ed764aa71fcb03746785b2f0aa08ce6e5c98f324bd7c8da8feb6a81d38a

Observation 7b0b1d20-208b-4754-8474-657d6830b089 · outbound

This paper cites Sigmoid loss for language image pre-training.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Sigmoid loss for language image pre-training

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.949397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e9ea6bb7f62f4e3ac35b4365513a2b285fb5b6e1667979befb87cb11e629ab8f

Observation 0a6dd454-d857-412b-8905-d434933e510d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.248894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:5b48af4a764c1be2ee207fea90cbc8e56f04a750eb874d2e9f95219fd893383f

Observation a91f9f70-303e-4390-a230-56c7157b4ba8 · outbound

This paper cites Vision trans- formers for dense prediction.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Vision trans- formers for dense prediction

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.017860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:5b8b5cbda0926f78c76601204310751d29dec966ba035ff95804da0f6ddd6857

Observation 38067fb4-b886-4c5c-8678-4df8ad88fafe · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.008267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:21e0a952f1776dcd318ac6fcabbe1b7bd5e821f8cd38d6a1bc1de6e8d0a778a5

Observation 0b80b7f0-6563-4815-853a-8376770dd978 · outbound

This paper cites Decoupled weight decay regularization.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Decoupled weight decay regularization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.002272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:1ff5a8348dc999bf6e0cfc28035be8ecdcc7318dc0e97672bb7708413b6c0d6a

Observation bb11bb1a-f882-465e-9f6d-256642ab8e5f · outbound

This paper cites In our real- robot experiments, we useK= 201atoms over[−0.1,1.1].

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies In our real- robot experiments, we useK= 201atoms over[−0.1,1.1]

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.025765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:bd064df831da76c1113d54ded88199fb765ad031b014074ee9a06e52ad025127

Observation fdb24c82-a66f-421d-9ad7-90692f548c5f · outbound

This paper cites an unresolved cited work.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-07-06T13:42:37.021848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:8dc073855f85084b8c8988b3db08e915272389103286f08cc070b8fd2d469a3a

Observation 1492b493-37db-4baf-b0a7-485dec5690ce · outbound

This paper cites an unresolved cited work.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-07-06T13:42:36.951626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:51895308836b2f796228a5c0d407b6c1d16c6e3a09c81d792b1eb27338b85c28

Observation 954c73f3-0eb1-4fad-ab9d-325d18f39c9b · outbound

This paper cites Demonstrations are successful trajectories, rollouts contain both successes and failures, and play data is treated as unsuccessful exploratory data.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Demonstrations are successful trajectories, rollouts contain both successes and failures, and play data is treated as unsuccessful exploratory data

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.940526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:dbdf7cb854e4c941a8e81a885e6d955039c55ca1573ff084b32e52303b07977e

Observation 58af25cc-d75f-4e8c-b548-e1ff60d54e12 · outbound

This paper cites The policy is optimized with AdamW [63] using a base learning rate of2×10 −5 and a cosine decay schedule.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The policy is optimized with AdamW [63] using a base learning rate of2×10 −5 and a cosine decay schedule

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.995805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:0465e842940e27da5a6644ae68705a0a057ec5ee38e73a142d80bc6b5c07069a

Observation c18fac3f-1573-45d4-be3e-c108997614bd · outbound

This paper cites an unresolved cited work.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-07-06T13:42:36.956522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2ff83aff6dd4c43102471c953398ba887b8e017c1a2514926739b325ebe41c44

Observation 396975de-941d-452b-8b0c-deb512f0c24a · outbound

This paper cites The model is trained with a flow-matching loss, where the interpolated noisy actiona w is defined in Eq.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The model is trained with a flow-matching loss, where the interpolated noisy actiona w is defined in Eq

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-07-06T13:42:36.961158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b00937b4d0a8f586db4a1f1aa1417d1595bb437ff6cd41c223de2289a9793aaa

Observation 18f88483-d5e8-43b7-b1f3-9a39d2806629 · outbound

This paper cites The comparison isolates the Robot 1 Robot 2.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The comparison isolates the Robot 1 Robot 2

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.991664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ca1d1dbc14069f7ebef9c5a72378d24a306f55430ae6b327993712ab56bc9cfd

Observation 595baa96-96fe-431c-bcb7-16445fcc92d6 · outbound

This paper cites 9 vi- sualizes the predicted value distributions for the same episodes shown in Fig.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies 9 vi- sualizes the predicted value distributions for the same episodes shown in Fig

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.954173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:a5e482e6b81a01a0f0b3bd2625ab169ef471b8489bd651c8b2e283f7a453ee87

Observation 09c10d04-97e4-410b-9d6e-efb8489b620d · outbound

This paper cites (i) Object-storage uploads commit atomically (read- ers see either the fully-uploaded payload or no object) and are retried until persisted.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (i) Object-storage uploads commit atomically (read- ers see either the fully-uploaded payload or no object) and are retried until persisted

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.982931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:16b1c918a03d2ddfd42d1b604cb2e9348bbdb1956d525ddf9f622f6338df50e6

Observation d87dbcdb-4e25-47a0-aec3-82a14d10eb4a · outbound

This paper cites Table VI reports both on the same 8-hour, 16-actor run as the End-to-End Reliability subsection above.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Table VI reports both on the same 8-hour, 16-actor run as the End-to-End Reliability subsection above

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.958802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d0fa15b7e231ad77fbc3c2b9636e5cbfd4630f1a94cd55195afb499011ea926d

Pith citing papers

Observation 6de04dfe-12cf-4fbe-8d39-258337ca15b7 · inbound

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning cites this paper.

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:58:02.685056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-27T09:46:59.746745Z digest=sha256:a00b1a36a14a092bae431b385376b895319859e195434163293955a46c4ee858

Observation 42d737be-edc7-42de-a44c-5013aaaf9047 · inbound

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation cites this paper.

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:59:42.035183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T10:50:24.868343Z digest=sha256:c010a126a23117bcfbeab1d5d9665725accb464debf789c57c625afd9236f494

Observation 30afc46d-77f9-40f2-8f24-ea6c22132644 · inbound

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning cites this paper.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:31cac25aaebf704eb11c89e0e5a7efba6bd4b5b8a4c4d2cbedebc0110627250d