Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:55:33.276019Z
Paper Citation Record · LEDGER
As of 21 July 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.03021.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:55:33.276019Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a5289c82-b752-4455-ba86-ca4d69de0dcf · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 1
Source-reported events for the cited work
Unavailable: named source frontier unavailable.
Observation 6138a529-2a89-436b-b052-7b601b501ef0 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 2
Source-reported events for the cited work
Unavailable: named source frontier unavailable.
Observation 6865d765-3e14-4e4c-ac25-15bbf44007fa · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Learning to Reason under Off-Policy Guidance
Reference 3
Source-reported events for the cited work
Unavailable: named source frontier unavailable.
Observation b467ea7f-5194-4989-a937-87073731b48d · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: named source frontier unavailable.
Observation 3cf3b544-01bf-4f2f-b598-1c80e09bdc7c · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 5
Source-reported events for the cited work
Unavailable: named source frontier unavailable.
Observation 0331e4cf-4cf3-454b-8d03-e71e6c1da721 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes” as a measure of the similarity between the two candi- date solutions. As shown by “HDPO (LLM-Div)
Reference 6
Source-reported events for the cited work
Unavailable: named source frontier unavailable.
Observation 92127c55-c70b-4017-b83b-83b0d2a6cadc · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12fcd6b0-9292-4ae0-b0a4-1c3ce572e676 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning ex- plore–evaluate–select
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8afd8b-3c22-4c2d-b26a-6d3b4782a808 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning It should only elaborate on the high-level strategies and concepts, without going into specific calculations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8484ce-4d34-4fbd-8f99-ea4358780e9c · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0722538-8053-4b5b-b6a5-23c426423162 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95241bac-5176-401c-8f33-36c21c0fa498 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning [1]”, “[2]
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fecd30-9f69-4150-9725-fcd524fd9603 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71784dfd-f297-4844-a6c2-b83b9754e482 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes", otherwise output
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87664f36-0060-4f7b-b371-5f53abe0e057 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c7ebac-94cd-4345-a996-9c63e6c097b3 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abacaf87-ad57-44f0-afa9-b0b5b76748c9 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning We need to find the values ofa,b, andcthat maximize|a|+|b|+|c|while satisfying these constraints
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca748070-bae6-45d3-89bc-0026455624cc · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Then express a, b, c in terms of these values and use linear programming or symmetry arguments to maximize |a| + |b| + |c|
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c50f582-1be0-47d3-8b2c-b8affaacb8a1 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Apply the method of Lagrange multipliers to maximize the linear functional |a| + |b| + |c| subject to the quadratic constraint
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a72dcb-d871-4c8d-942c-993140c365fe · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this may miss the global maximum if the optimal polynomial is not symmetric or has non-zero a and b
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd6b323-e61e-4f1b-aad7-04ed6e6a14da · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Scale and shift the Chebyshev polynomial to satisfy the bound|P(x)| ≤1and compute the coefficients to find the maximum of |a| + |b| + |c|
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06aea44a-95a3-4876-8b8e-324c355a9784 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db848297-b0a0-4389-a502-48367449eb9e · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning - PointPis 4 units away from the circle, so the distance fromPto the centerOis6 + 4 = 10
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea92010a-ee05-4c91-b1eb-61b6c1741a0d · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning But sincePis 4 units away from the circle and AB is parallel to ← →OP, the perpendicular distance fromOto ABmust be 4 (as 8 would place AB outside the circle)
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a49c998f-77bf-426b-8104-e315f03ea951 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning - The chord ABis parallel to the x-axis and 2 units below the x-axis (since the distance from Oto ABis 4)
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e5b615-a66f-47a8-95a3-63c9337fd251 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb498f7e-d7a3-41f3-bc09-7119a2cc1a85 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning 19 Case 2 (Generation Model: Qwen2.5-Math-7B-HDPO) Question: CircleOhas radius 6
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4b1468-b2e5-428d-add9-19958630381e · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning SinceABis parallel toOP and the distance between them is 2, the perpendicular distance fromOtoABis either4 + 2 = 6 or4−2 = 2
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a27ee0-65a8-486e-8b41-3e48c3cc7f62 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1bcde0-f444-4255-906d-6ae29d70119e · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Since chordABis parallel to ← →OP, it is horizontal, and the distance betweenAB and ← →OPis 2, soABis either aty= 2ory=−2
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04f6e3f-c5f9-456a-b7c2-60bc57b60bfa · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252f6b2d-5468-478d-8941-fa2f24bf7eb3 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Then apply the sum of cosine series formula for angles in arithmetic sequence, simplifying the resulting expression using symmetry and periodicity of the cosine function
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f4b924-04b7-4722-aad8-9cd5106fbf03 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this approach lacks precision and relies on approximation, making it unsuitable for exact computation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa14e551-b1c2-478d-8bb3-6488dd20bb50 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e0e174-2b13-4aaa-a159-b94a4b861385 · outbound
Hint-Guided Diversified Policy Optimization for LLM Reasoning </Candidate Solutions> <selected>[1]</selected> <thinking> We are given the sum: sin2 4◦ + sin2 8◦ + sin2 12◦ +· · ·+ sin 2 176◦ This is a sum ofsin 2 θforθ= 4k ◦ wherek= 1,2,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.