Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T22:43:50.317854Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2605.10325.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T22:43:50.317854Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 972a20b1-7617-496f-a8cf-7bf7c87c0396 · outbound
Verifiable Process Rewards for Agentic Reasoning FireAct: Toward Language Agent Fine-tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 28d68363-e075-400d-88e4-1dd15db9bf91 · outbound
Verifiable Process Rewards for Agentic Reasoning Group-in-Group Policy Optimization for LLM Agent Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation e546cefc-06e5-4d26-b60f-afbd517f66e2 · outbound
Verifiable Process Rewards for Agentic Reasoning CRITIC: Large language models can self-correct with tool-interactive critiquing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 0578d6c8-30b5-4333-be70-fa8f205b23bd · outbound
Verifiable Process Rewards for Agentic Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 9bd9ee2b-67bb-489d-99bc-25fa4d7eb2de · outbound
Verifiable Process Rewards for Agentic Reasoning Large language models cannot self-correct reasoning yet
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation f3b3a2cc-5b4c-4840-a644-0bbb4c94a836 · outbound
Verifiable Process Rewards for Agentic Reasoning OpenAI o1 System Card
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation ad3f1f7f-e76b-48ac-a87e-108d179f4c3e · outbound
Verifiable Process Rewards for Agentic Reasoning SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c9c4ef3c-4e1f-4886-abe7-33530a1846b1 · outbound
Verifiable Process Rewards for Agentic Reasoning Littman, and Anthony R
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1ac6339c-a008-457a-a54f-8270183a8e05 · outbound
Verifiable Process Rewards for Agentic Reasoning VinePPO: Refining credit assignment in RL training of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1f334c22-9fa2-49d4-8436-06d1cf87f6b3 · outbound
Verifiable Process Rewards for Agentic Reasoning Adam: A Method for Stochastic Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 48607731-3857-44a2-aad5-3f6dc362677d · outbound
Verifiable Process Rewards for Agentic Reasoning Bandit based monte-carlo planning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 7499655f-d83c-445a-80a0-b00424c110d9 · outbound
Verifiable Process Rewards for Agentic Reasoning Gonzalez, Hao Zhang, and Ion Stoica
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 37361511-ef60-42b5-9dc2-65234ef55cdc · outbound
Verifiable Process Rewards for Agentic Reasoning OpenSpiel: A Framework for Reinforcement Learning in Games
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1b23542e-f6f7-4d4d-a8a2-95aa2dc6d2dc · outbound
Verifiable Process Rewards for Agentic Reasoning Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 99e1009a-9757-4cf1-8b13-537693c5b201 · outbound
Verifiable Process Rewards for Agentic Reasoning Let’s verify step by step
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation ce3f39f1-ec20-45d7-adb8-d4d2b9e11748 · outbound
Verifiable Process Rewards for Agentic Reasoning Agentbench: Evaluating LLMs as agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation ff74810b-a1f4-406c-9f32-d6be32d99f45 · outbound
Verifiable Process Rewards for Agentic Reasoning Gem: A gym for agentic llms.arXiv preprint arXiv:2510.01051
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c2c685dc-5c88-407c-ab7e-6141e561e43f · outbound
Verifiable Process Rewards for Agentic Reasoning Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation cc81527e-5ff7-467d-b742-b7951f43d538 · outbound
Verifiable Process Rewards for Agentic Reasoning Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation b6cea56b-7668-4f1c-aa77-e1fe8d99785a · outbound
Verifiable Process Rewards for Agentic Reasoning Code Llama: Open Foundation Models for Code
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 860a9b7f-1d3e-44d1-b925-b0d02227742c · outbound
Verifiable Process Rewards for Agentic Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 27722fbe-39b1-419d-ab33-6360a2d6b010 · outbound
Verifiable Process Rewards for Agentic Reasoning Reflex- ion: language agents with verbal reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation a1c5e64e-bf48-4087-b923-ba3600206eae · outbound
Verifiable Process Rewards for Agentic Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d696dc85-9849-4a0f-8910-d0fdaa5a94d7 · outbound
Verifiable Process Rewards for Agentic Reasoning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 12e2a025-5abc-47cf-963b-0a10cc109017 · outbound
Verifiable Process Rewards for Agentic Reasoning EvalScope: Evaluation framework for large models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation de50c423-a368-4788-aa9d-68ae1ffd5e88 · outbound
Verifiable Process Rewards for Agentic Reasoning Solving math word problems with process- and outcome-based feedback
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c6b170ee-9ce5-4b2e-b329-20517261f681 · outbound
Verifiable Process Rewards for Agentic Reasoning A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation bc9d67fa-b83c-4b52-8db2-ef3edbb56927 · outbound
Verifiable Process Rewards for Agentic Reasoning Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation b1f48d73-be66-4a87-b2f0-3d1e3bde7565 · outbound
Verifiable Process Rewards for Agentic Reasoning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation a978a41f-9ca9-4c56-a665-4e8af130be24 · outbound
Verifiable Process Rewards for Agentic Reasoning what it can create, it may not understand
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d88c47ad-f4b6-4297-b7b0-52748047c7c9 · outbound
Verifiable Process Rewards for Agentic Reasoning The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c9137a1d-e760-4b91-9f1e-fb266d219856 · outbound
Verifiable Process Rewards for Agentic Reasoning Vs-bench: Evaluating vlms for strategic reasoning and decision-making in multi- agent environments.coming soon
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 76c406f7-fe8a-4efc-9a48-8fb3507f2d5b · outbound
Verifiable Process Rewards for Agentic Reasoning Qwen3 Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1f33ab50-0732-4a5f-a6af-3709d3937001 · outbound
Verifiable Process Rewards for Agentic Reasoning Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 65915e37-c966-4622-8875-90ebd5734ff7 · outbound
Verifiable Process Rewards for Agentic Reasoning Tree of thoughts: Deliberate problem solving with large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 13a0a50e-9b59-44bd-aa7a-8a8ec5a38efd · outbound
Verifiable Process Rewards for Agentic Reasoning React: Synergizing reasoning and acting in language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation eb0901ef-55b0-4aec-b12f-e206e6b4f29a · outbound
Verifiable Process Rewards for Agentic Reasoning OVM, outcome-supervised value models for planning in mathematical reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c78e3567-40b2-4521-ba82-8d3e82868f0c · outbound
Verifiable Process Rewards for Agentic Reasoning Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 2ae4d594-3b51-4d7f-9555-2932e38cecf3 · outbound
Verifiable Process Rewards for Agentic Reasoning Language agent tree search unifies reasoning, acting, and planning in language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation f46590e2-3719-4ea3-b2a0-b07c4cf6a2f9 · outbound
Verifiable Process Rewards for Agentic Reasoning The grid is 0 - indexed , where (0 ,0) is the top - left corner and (2 ,2) is the bottom - right corner
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 0ce9e46a-032d-4770-a71c-d20f812bb9df · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 91b1d822-857e-4bb8-a757-1fe97e4ac311 · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 299c8de3-d4dd-48d9-87ac-adabb2f8f2bb · outbound
Verifiable Process Rewards for Agentic Reasoning PLAYER I N F O R M A T I O N
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 22705732-7545-4157-8efe-22b2bd4a6bd1 · outbound
Verifiable Process Rewards for Agentic Reasoning You are c om pe ti ng with another player c o n t r o l l i n g the mark O
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 87153c57-244c-423a-9fa0-19b30cbc9c60 · outbound
Verifiable Process Rewards for Agentic Reasoning The game state d e m o n s t r a t e s the current board with a three - line text grid , where ’X ’ and ’O ’ are the marks of the two players , and ’
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 6bc4715f-fc38-4962-9f11-ce678fc76882 · outbound
Verifiable Process Rewards for Agentic Reasoning Rows and columns are 1 - indexed (1 to 9)
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation ffec39ca-ed84-4bff-b097-bd16002841e5 · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 671aafb7-5934-460d-bf5b-b6ce5a2b3f90 · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 0cd13fc0-f01c-4fea-bf97-38f37693580c · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 5b8359b3-3656-473a-81e2-26cb0d330c6a · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation a25f6b78-159d-4aeb-b3fc-496acc6eb1c2 · outbound
Verifiable Process Rewards for Agentic Reasoning PLAYER I N F O R M A T I O N
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 9c06e445-60bd-4d07-b323-40762c49d60b · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 46264534-104e-4cd0-9246-5be36e95240f · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 506bcb24-3c28-4919-9495-b2670963201c · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation deca4e3b-fde5-4b4f-b026-2dac68379439 · outbound
Verifiable Process Rewards for Agentic Reasoning The grid contains exactly 5 hidden mines
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 33fa1e48-eb83-4eeb-ae84-c404ca20dc7c · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1f36f8ef-83d9-429a-a01c-ddceb9e8be29 · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 62122df2-ee61-45eb-beb4-c7c71ea240ca · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 090c2d4e-10bd-4110-b158-f7634604edd6 · outbound
Verifiable Process Rewards for Agentic Reasoning PLAYER I N F O R M A T I O N
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation c4f69ac0-3796-4f9a-af88-970be55f2783 · outbound
Verifiable Process Rewards for Agentic Reasoning ’ r e p r e s e n t s an u n r e v e a l e d cell
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d636f102-1bb3-4f90-b4dd-ee92339b4063 · outbound
Verifiable Process Rewards for Agentic Reasoning Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 7cc632f8-63b5-47e0-9d54-efa62b7f0c06 · outbound
Verifiable Process Rewards for Agentic Reasoning The ’ flag ’ command acts as a toggle : play it on an un fl agg ed cell to place a flag , or on a flagged cell to remove it
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
No inbound Pith citation observations are available.