Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:42:37.065795Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 5 of 5 outbound references and 6 inbound Pith citation observations for arXiv:2604.07165.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:42:37.065795Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T16:03:26.866407Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T20:47:23.391702Z
5 of 5 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a2136d82-fca9-4b9c-845f-e4684b6422a3 · outbound
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization Group-in-Group Policy Optimization for LLM Agent Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation dce55c31-fb94-4643-8380-ddd97fda185c · outbound
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization arXiv preprint arXiv:2509.09284 , year=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation cc0dfc80-e603-459e-aafc-9d4590d40378 · outbound
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 15473c44-3a79-4937-baf1-6dd658ce8f99 · outbound
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 7e08366f-f6c2-460f-bb88-9c68f0ec2935 · outbound
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization assembly required
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 11925afe-9ba5-4e32-83f5-4ea16ec3fd39 · inbound
NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 1c8b7a9d-e82c-4743-9563-556c4959924d · inbound
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 2d916b6f-2bd0-4a88-b0fa-d45a06792c15 · inbound
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation d900fad7-4b4f-40a2-9b43-c2258b6e6922 · inbound
A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 6241198f-c8ad-4dd2-b082-203b58cf95ec · inbound
Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.
Observation 3d7b86d3-f1ce-4a28-b794-d8dd0834cbe6 · inbound
Customer-Agent: Overcoming Context Limitations in Ultra-Long Shopping Trajectories via Tool-Augmented Agents and RLVR Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.