Pith. sign in

Paper Citation Record · LEDGER

Verifiable Process Rewards for Agentic Reasoning

As of 22 July 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2605.10325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10325 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:43:50.317854Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact14
  • verified fuzzy35
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 972a20b1-7617-496f-a8cf-7bf7c87c0396 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

Verifiable Process Rewards for Agentic Reasoning FireAct: Toward Language Agent Fine-tuning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:07.256818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:8ce9563707ba38dfb7541345e50f1dc8663df3648f2652ae86dd4929f2772d60

Observation 28d68363-e075-400d-88e4-1dd15db9bf91 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Verifiable Process Rewards for Agentic Reasoning Group-in-Group Policy Optimization for LLM Agent Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.246790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:e9b7eddd4abe6bda33d8931486cc126ae6da64df5c3a824dc2b81f9e78822732

Observation e546cefc-06e5-4d26-b60f-afbd517f66e2 · outbound

This paper cites CRITIC: Large language models can self-correct with tool-interactive critiquing.

Verifiable Process Rewards for Agentic Reasoning CRITIC: Large language models can self-correct with tool-interactive critiquing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.127328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f5fc4536cce82c8b9e4c7c62c24cb2f86833223128ec9614050c48883b8b1ba0

Observation 0578d6c8-30b5-4333-be70-fa8f205b23bd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Verifiable Process Rewards for Agentic Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.269804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:0758755d86329231e1ab6bd03f8bcc0171d4461068743ee176706b240408b4f9

Observation 9bd9ee2b-67bb-489d-99bc-25fa4d7eb2de · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Verifiable Process Rewards for Agentic Reasoning Large language models cannot self-correct reasoning yet

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.115481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:200be91cdbab2f928ce943311f8469454081c9beeee85a5550574e154a69e311

Observation f3b3a2cc-5b4c-4840-a644-0bbb4c94a836 · outbound

This paper cites OpenAI o1 System Card.

Verifiable Process Rewards for Agentic Reasoning OpenAI o1 System Card

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.243977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:2386fa073173eaf2f655a3499b7d5965b539a338389fef9b4c9e72cb0d46f982

Observation ad3f1f7f-e76b-48ac-a87e-108d179f4c3e · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations.

Verifiable Process Rewards for Agentic Reasoning SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.119571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:652528ae8abbbad11f7850870128c77a6d62d3375c16d012c7217ba623c010bd

Observation c9c4ef3c-4e1f-4886-abe7-33530a1846b1 · outbound

This paper cites Littman, and Anthony R.

Verifiable Process Rewards for Agentic Reasoning Littman, and Anthony R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.219129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:a9006b95c99c484c11101419b2ed288e274b6ac48327db373897ef80e8c75c7a

Observation 1ac6339c-a008-457a-a54f-8270183a8e05 · outbound

This paper cites VinePPO: Refining credit assignment in RL training of LLMs.

Verifiable Process Rewards for Agentic Reasoning VinePPO: Refining credit assignment in RL training of LLMs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.108491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:03d8c5d2ed190ac856da5b47c4f4333c7676d9a9d92d757ed098041baace52e1

Observation 1f334c22-9fa2-49d4-8436-06d1cf87f6b3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Verifiable Process Rewards for Agentic Reasoning Adam: A Method for Stochastic Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.263078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f644a91ce4bb8ca6e0ddefff6a65061b9fa9011b79cb871f30d97e445d084d2f

Observation 48607731-3857-44a2-aad5-3f6dc362677d · outbound

This paper cites Bandit based monte-carlo planning.

Verifiable Process Rewards for Agentic Reasoning Bandit based monte-carlo planning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.104264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:09facd4d1b19b402021509ee4de18e7418e309983cb095522b5833df6328670d

Observation 7499655f-d83c-445a-80a0-b00424c110d9 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Verifiable Process Rewards for Agentic Reasoning Gonzalez, Hao Zhang, and Ion Stoica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.137430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:b8e8a1940e3e98921cbd9ed7c15f63c12bc648edad9b808252820e9a2ea733e7

Observation 37361511-ef60-42b5-9dc2-65234ef55cdc · outbound

This paper cites OpenSpiel: A Framework for Reinforcement Learning in Games.

Verifiable Process Rewards for Agentic Reasoning OpenSpiel: A Framework for Reinforcement Learning in Games

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:07.249837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:a7f883de67beaeb5cffe17a357700ef702b9792516445252a199548763a13fca

Observation 1b23542e-f6f7-4d4d-a8a2-95aa2dc6d2dc · outbound

This paper cites Coderl: Mastering code generation through pretrained models and deep reinforcement learning.

Verifiable Process Rewards for Agentic Reasoning Coderl: Mastering code generation through pretrained models and deep reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.145041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:a498cba39561237b11905bfafbef26c0e95d3d6ae21912e97b956585c91dfbd5

Observation 99e1009a-9757-4cf1-8b13-537693c5b201 · outbound

This paper cites Let’s verify step by step.

Verifiable Process Rewards for Agentic Reasoning Let’s verify step by step

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.161180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:e2ae4ad10513261b9c0b299c963bd381acddded6587fbecc930631cecc43ad8e

Observation ce3f39f1-ec20-45d7-adb8-d4d2b9e11748 · outbound

This paper cites Agentbench: Evaluating LLMs as agents.

Verifiable Process Rewards for Agentic Reasoning Agentbench: Evaluating LLMs as agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.176836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:80761f866a721b22514ace4a97cd1e6fb65f929f7b140f9ae1c581b233ab97e4

Observation ff74810b-a1f4-406c-9f32-d6be32d99f45 · outbound

This paper cites Gem: A gym for agentic llms.arXiv preprint arXiv:2510.01051.

Verifiable Process Rewards for Agentic Reasoning Gem: A gym for agentic llms.arXiv preprint arXiv:2510.01051

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:07.253248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:4eb272d6da822b6781dfed4d44b9d815b26aaa301183ff7835b67faadae33f67

Observation c2c685dc-5c88-407c-ab7e-6141e561e43f · outbound

This paper cites Training language models to follow instructions with human feedback.

Verifiable Process Rewards for Agentic Reasoning Training language models to follow instructions with human feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.179603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:c2fd3771b8527700313b686d15ec40f8478c7d7b8dd7834b15efe30105b9e2e5

Observation cc81527e-5ff7-467d-b742-b7951f43d538 · outbound

This paper cites Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning.

Verifiable Process Rewards for Agentic Reasoning Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.142765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:7793191d6fb50cca93a29199eeb0f5dcf8b19799e697a3339bea456b9679fa91

Observation b6cea56b-7668-4f1c-aa77-e1fe8d99785a · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Verifiable Process Rewards for Agentic Reasoning Code Llama: Open Foundation Models for Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.259553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:faeeff4f5ac1d57c4e829b1857b34e4775156dc5fbec6c9f61e492efc5123b59

Observation 860a9b7f-1d3e-44d1-b925-b0d02227742c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Verifiable Process Rewards for Agentic Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.273054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:07d453d397bc350cc10bcdc241523cfa64f1a0f1edb6aa298876066f49c4653a

Observation 27722fbe-39b1-419d-ab33-6360a2d6b010 · outbound

This paper cites Reflex- ion: language agents with verbal reinforcement learning.

Verifiable Process Rewards for Agentic Reasoning Reflex- ion: language agents with verbal reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.172414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:a4cb8b6a16f2e223ae3f19a2cd08a1bbbed7f22c7194f66ac3b931fd1ae7936f

Observation a1c5e64e-bf48-4087-b923-ba3600206eae · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Verifiable Process Rewards for Agentic Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.236895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:7d31fc744a6b0c56365c93fdc3cd840dcad930151f98e61621394f55b5ca3c08

Observation d696dc85-9849-4a0f-8910-d0fdaa5a94d7 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Verifiable Process Rewards for Agentic Reasoning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.234587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:277d5b1b9f2e1534f4483bc659f2d40ff9470ba86cc5868b223dd33c4dabba6c

Observation 12e2a025-5abc-47cf-963b-0a10cc109017 · outbound

This paper cites EvalScope: Evaluation framework for large models.

Verifiable Process Rewards for Agentic Reasoning EvalScope: Evaluation framework for large models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.170042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f4140dcbc28b967f0df3deb4298bfd653c777b54aae897b577a7b071bfe2a9a0

Observation de50c423-a368-4788-aa9d-68ae1ffd5e88 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Verifiable Process Rewards for Agentic Reasoning Solving math word problems with process- and outcome-based feedback

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.239253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:6bf35f429370ef442ca3f602aeacdea13aa8a4d92c14f03e1c984a1137a27bfb

Observation c6b170ee-9ce5-4b2e-b329-20517261f681 · outbound

This paper cites A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345.

Verifiable Process Rewards for Agentic Reasoning A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.174679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:da985ea678e72a7cd7dc8d3c81d95ec5a71f157e3e610b3bf461ac80a681b846

Observation bc9d67fa-b83c-4b52-8db2-ef3edbb56927 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Verifiable Process Rewards for Agentic Reasoning Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.182578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:affdc4692cc769d58c99920ea4cb4501996720d277b4171f82b8668b66bf2650

Observation b1f48d73-be66-4a87-b2f0-3d1e3bde7565 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

Verifiable Process Rewards for Agentic Reasoning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:45:07.266701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:1e65edf1d103ba8e31602632448a624926ac4312a61560bbb8a7ce796c78da9a

Observation a978a41f-9ca9-4c56-a665-4e8af130be24 · outbound

This paper cites what it can create, it may not understand.

Verifiable Process Rewards for Agentic Reasoning what it can create, it may not understand

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.188796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:e27bfc427c1f21051af5cfdbac0769841e411a41b781a0db68b6f9ee217b3241

Observation d88c47ad-f4b6-4297-b7b0-52748047c7c9 · outbound

This paper cites The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101.

Verifiable Process Rewards for Agentic Reasoning The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.194378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:32d15aacfc012ec536074bae2924046bc40f66c9a867154e11d16c688571f4af

Observation c9137a1d-e760-4b91-9f1e-fb266d219856 · outbound

This paper cites Vs-bench: Evaluating vlms for strategic reasoning and decision-making in multi- agent environments.coming soon.

Verifiable Process Rewards for Agentic Reasoning Vs-bench: Evaluating vlms for strategic reasoning and decision-making in multi- agent environments.coming soon

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.192083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:82895ce35ac13fd54e06728e9176af94f3819ed9d49ca6f924555774a7d898a4

Observation 76c406f7-fe8a-4efc-9a48-8fb3507f2d5b · outbound

This paper cites Qwen3 Technical Report.

Verifiable Process Rewards for Agentic Reasoning Qwen3 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:07.241651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:89dc8088da216b47c3115ce94ce9132d2017ffba84ca1c9c692c3cb6778faa9e

Observation 1f33ab50-0732-4a5f-a6af-3709d3937001 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757.

Verifiable Process Rewards for Agentic Reasoning Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.158958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:2889218f87be5258b4aee0b79cfbcf5f7517588901f2aa46e610c0f19020a4cb

Observation 65915e37-c966-4622-8875-90ebd5734ff7 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Verifiable Process Rewards for Agentic Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.163459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:71ae69b8f8c74f368a988c2daab6b658213eb0eaa4dc21991d1f9a5eae0095b2

Observation 13a0a50e-9b59-44bd-aa7a-8a8ec5a38efd · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Verifiable Process Rewards for Agentic Reasoning React: Synergizing reasoning and acting in language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.165600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f8f973c9c4549a4fe297a8f141272ef0695a748e289aed4439dece22390c25c7

Observation eb0901ef-55b0-4aec-b12f-e206e6b4f29a · outbound

This paper cites OVM, outcome-supervised value models for planning in mathematical reasoning.

Verifiable Process Rewards for Agentic Reasoning OVM, outcome-supervised value models for planning in mathematical reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.153866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:afcc1a88993351a0cedfce1986dee759dd640c8ba7f84001c9dd8d86aaac759c

Observation c78e3567-40b2-4521-ba82-8d3e82868f0c · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Verifiable Process Rewards for Agentic Reasoning Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.123499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:5f64c4634f7575fd2fc570fc5d4b3e59fc5f2eaef378e71b29cc4e8c032d8d40

Observation 2ae4d594-3b51-4d7f-9555-2932e38cecf3 · outbound

This paper cites Language agent tree search unifies reasoning, acting, and planning in language models.

Verifiable Process Rewards for Agentic Reasoning Language agent tree search unifies reasoning, acting, and planning in language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.140444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:114b03bca1fc1ca781b0bb5c9f9da610b140e2c46ede8385acbff2338e65f2c7

Observation f46590e2-3719-4ea3-b2a0-b07c4cf6a2f9 · outbound

This paper cites The grid is 0 - indexed , where (0 ,0) is the top - left corner and (2 ,2) is the bottom - right corner.

Verifiable Process Rewards for Agentic Reasoning The grid is 0 - indexed , where (0 ,0) is the top - left corner and (2 ,2) is the bottom - right corner

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.151547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:b1c27de91d3123164ef804004ff80708a639cb3da0bfe7a549c4bbe2557e12fe

Observation 0ce9e46a-032d-4770-a71c-d20f812bb9df · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.130514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:c33eb971b95b3394c381b79c9518713dae65fbf967a071c6a9deba0d24bca0cd

Observation 91b1d822-857e-4bb8-a757-1fe97e4ac311 · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.111778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:9c71d210a70f75efe7e6934fea6f28c56a7b3a4268790840fe88b36bb344613c

Observation 299c8de3-d4dd-48d9-87ac-adabb2f8f2bb · outbound

This paper cites PLAYER I N F O R M A T I O N.

Verifiable Process Rewards for Agentic Reasoning PLAYER I N F O R M A T I O N

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.156290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:e73100c6365d04c22a01248e7fb4d2e6ae597eacdd6ac5a0bed70e5864fd5d5e

Observation 22705732-7545-4157-8efe-22b2bd4a6bd1 · outbound

This paper cites You are c om pe ti ng with another player c o n t r o l l i n g the mark O.

Verifiable Process Rewards for Agentic Reasoning You are c om pe ti ng with another player c o n t r o l l i n g the mark O

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.222635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:fb47c9aebaa400e446e4e585c10241282d2af60bdc040413fb94ca5b30c7e178

Observation 87153c57-244c-423a-9fa0-19b30cbc9c60 · outbound

This paper cites The game state d e m o n s t r a t e s the current board with a three - line text grid , where ’X ’ and ’O ’ are the marks of the two players , and ’.

Verifiable Process Rewards for Agentic Reasoning The game state d e m o n s t r a t e s the current board with a three - line text grid , where ’X ’ and ’O ’ are the marks of the two players , and ’

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.147396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:a4e4990bf520e57b428ea20d9cce77a80938068ecbbf43b0a6cc5980f25221de

Observation 6bc4715f-fc38-4962-9f11-ce678fc76882 · outbound

This paper cites Rows and columns are 1 - indexed (1 to 9).

Verifiable Process Rewards for Agentic Reasoning Rows and columns are 1 - indexed (1 to 9)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.133760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:094462269655de1ba949f87787906e7ce1057a6fad48ce03c97666b2077f3477

Observation ffec39ca-ed84-4bff-b097-bd16002841e5 · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.149370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:4803db176dafdfb9334624272a4da2b4e313396c72c6ca75b77ddda02ccfc9fb

Observation 671aafb7-5934-460d-bf5b-b6ce5a2b3f90 · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.167691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f2169dc908e98087282771c4a790453aa4fe02e5f60974b86fd5be5329fcf83e

Observation 0cd13fc0-f01c-4fea-bf97-38f37693580c · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.185373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:98688657e88ae680ce68899165858b1650370bab330e3d6c960abe6a85e00fb4

Observation 5b8359b3-3656-473a-81e2-26cb0d330c6a · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.216936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:c661b3bc434ae58fa0c697422ceac42b719fa526ec81e79f1d19df9a9cded30f

Observation a25f6b78-159d-4aeb-b3fc-496acc6eb1c2 · outbound

This paper cites PLAYER I N F O R M A T I O N.

Verifiable Process Rewards for Agentic Reasoning PLAYER I N F O R M A T I O N

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.220818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:4971b36f6fbc7f2081dd81cee463691be5593006bdcdf831429e456e5837735e

Observation 9c06e445-60bd-4d07-b323-40762c49d60b · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.224549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f24ae0212c50a15dcbdb80ddcef4814190c3a636fcd7e86bd09e61b2109bc64e

Observation 46264534-104e-4cd0-9246-5be36e95240f · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.207797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:5e55617768bee93d3477d2aa3ebab52b638297fd687377a8c83ba119d4935faf

Observation 506bcb24-3c28-4919-9495-b2670963201c · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.205754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:02522be055fe7e71fd3efe169b9f2d0dd483c72390f6c0daf6ee098a94db6446

Observation deca4e3b-fde5-4b4f-b026-2dac68379439 · outbound

This paper cites The grid contains exactly 5 hidden mines.

Verifiable Process Rewards for Agentic Reasoning The grid contains exactly 5 hidden mines

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.209812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:8748990295e06207292130a0c3ef53ff344d60a05087e5aa7f08448ce5008813

Observation 33fa1e48-eb83-4eeb-ae84-c404ca20dc7c · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.211987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:cde855ecb50c6327e6d6cd0c079358da469d052d2fb29ceae988663a1b1cc76f

Observation 1f36f8ef-83d9-429a-a01c-ddceb9e8be29 · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.203883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:c739c04338deaf0700a259bb768604da67767616c51d8cec5188ff029bbf4c5f

Observation 62122df2-ee61-45eb-beb4-c7c71ea240ca · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.198227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:9a39351bd8146f33f7e46b2193f2904fc19e2b396e3bd870aaeb8d1941f980f0

Observation 090c2d4e-10bd-4110-b158-f7634604edd6 · outbound

This paper cites PLAYER I N F O R M A T I O N.

Verifiable Process Rewards for Agentic Reasoning PLAYER I N F O R M A T I O N

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.196374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:fe2db46e8998054a6bf297328d91bacac6f1e6f88fe1ce82e9e4608ea5d61c40

Observation c4f69ac0-3796-4f9a-af88-970be55f2783 · outbound

This paper cites ’ r e p r e s e n t s an u n r e v e a l e d cell.

Verifiable Process Rewards for Agentic Reasoning ’ r e p r e s e n t s an u n r e v e a l e d cell

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.200141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:b2102a5ca1c8cc500daeb1a65f7dea34ef51a21b4968219f62705e4f0cb98f09

Observation d636f102-1bb3-4f90-b4dd-ee92339b4063 · outbound

This paper cites an unresolved cited work.

Verifiable Process Rewards for Agentic Reasoning Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-07-07T13:03:52.202028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:7bc1ae698a96fbf534404673d114d3b00f3d1a3240f48cc2874505e048b08ad4

Observation 7cc632f8-63b5-47e0-9d54-efa62b7f0c06 · outbound

This paper cites The ’ flag ’ command acts as a toggle : play it on an un fl agg ed cell to place a flag , or on a flagged cell to remove it.

Verifiable Process Rewards for Agentic Reasoning The ’ flag ’ command acts as a toggle : play it on an un fl agg ed cell to place a flag , or on a flagged cell to remove it

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T13:03:52.214159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T22:43:50.317854Z digest=sha256:f288cdf30053359e7130f64db33d62cc7fba811ec4f894afb5b1429327c9d77f

Pith citing papers

No inbound Pith citation observations are available.