Pith. sign in

Paper Citation Record · LEDGER

Hint-Guided Diversified Policy Optimization for LLM Reasoning

As of 21 July 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.03021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03021 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:55:33.276019Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved28
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a5289c82-b752-4455-ba86-ca4d69de0dcf · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Hint-Guided Diversified Policy Optimization for LLM Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:26:26.885472Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:5f42f235ca01fbfaf29946ac047dca5d1b272ce0db6cc135e26be79ee39e5bf8

Observation 6138a529-2a89-436b-b052-7b601b501ef0 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:26:26.899220Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:79c2c280723c85bc2a95d8c7303fdcbf444596ff8b0f970721034d602f07cd55

Observation 6865d765-3e14-4e4c-ac25-15bbf44007fa · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Learning to Reason under Off-Policy Guidance

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.888167Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:bb20e58dce56e6971570ac4bc889390b43e6736c99395623b19cb552a64c074c

Observation b467ea7f-5194-4989-a937-87073731b48d · outbound

This paper cites Qwen3 Technical Report.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Qwen3 Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:26:26.893490Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:4c5a105532fb514a49d307211243796a453996f5009ca34fe98c63840411bb47

Observation 3cf3b544-01bf-4f2f-b598-1c80e09bdc7c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Hint-Guided Diversified Policy Optimization for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:26:26.896086Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:e570208ddd0743a27708a084053f1c6858efb69aa655de80262ba2dda2905d54

Observation 0331e4cf-4cf3-454b-8d03-e71e6c1da721 · outbound

This paper cites Yes” as a measure of the similarity between the two candi- date solutions. As shown by “HDPO (LLM-Div).

Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes” as a measure of the similarity between the two candi- date solutions. As shown by “HDPO (LLM-Div)

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.890915Z

Source-reported events for the cited work

Unavailable: named source frontier unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:e4863b99d97a5e97e0feedc4419506fbc922c8921f466f7073638d554af1e1e7

Observation 92127c55-c70b-4017-b83b-83b0d2a6cadc · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:b5b2d87143a546963b632200daeffebcf16d6d7572f8e0a79f951249f5c59ac9

Observation 12fcd6b0-9292-4ae0-b0a4-1c3ce572e676 · outbound

This paper cites ex- plore–evaluate–select.

Hint-Guided Diversified Policy Optimization for LLM Reasoning ex- plore–evaluate–select

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:800d631ee9d8585840e620ea13b7c2f86b69a910f83bf597acb7730e7299ca77

Observation 9e8afd8b-3c22-4c2d-b26a-6d3b4782a808 · outbound

This paper cites It should only elaborate on the high-level strategies and concepts, without going into specific calculations.

Hint-Guided Diversified Policy Optimization for LLM Reasoning It should only elaborate on the high-level strategies and concepts, without going into specific calculations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:6e05af384a261c1c11e98a27ff1d0eb1e7a782a27b6c334fa54fdba6d4a3a8a2

Observation 3e8484ce-4d34-4fbd-8f99-ea4358780e9c · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:8c062b95bd0f9e1a3ed51f878fc0d3fa4963eb58ce2c58e0d5efe94f42db208f

Observation c0722538-8053-4b5b-b6a5-23c426423162 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:9fa9f1c6685d2f66e121e3976aab2d602bf417c45cda47986027fd08c1acde26

Observation 95241bac-5176-401c-8f33-36c21c0fa498 · outbound

This paper cites [1]”, “[2].

Hint-Guided Diversified Policy Optimization for LLM Reasoning [1]”, “[2]

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:8df1a97bf3efc45669a0e0d308e8e439c71c0536a78cd96128fd60b0b3f8a9d6

Observation e8fecd30-9f69-4150-9725-fcd524fd9603 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:2f25eda5511f9e9d84800504c3f9030187a3425bbb15ef61eeff38a849ed01ae

Observation 71784dfd-f297-4844-a6c2-b83b9754e482 · outbound

This paper cites Yes", otherwise output.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes", otherwise output

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:354a45740083fcd3c6101e8da1acd9ff3b489eacf753b4e00cf2c54cde52f23a

Observation 87664f36-0060-4f7b-b371-5f53abe0e057 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 15

Resolution
parse uncertain
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:e1428187fa8727c134179ee71511d6992ab1a2bce1d29f434ecc6cc51801abf0

Observation 47c7ebac-94cd-4345-a996-9c63e6c097b3 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:1c92a348037140c0fbafbf95d1b15f370bfef9b937249ef66582668c7a5bb76a

Observation abacaf87-ad57-44f0-afa9-b0b5b76748c9 · outbound

This paper cites We need to find the values ofa,b, andcthat maximize|a|+|b|+|c|while satisfying these constraints.

Hint-Guided Diversified Policy Optimization for LLM Reasoning We need to find the values ofa,b, andcthat maximize|a|+|b|+|c|while satisfying these constraints

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:ff4dfdbe88a2d6e2f8f510091adcd3de91b3fdfe9021baf98727815e48ad2e19

Observation ca748070-bae6-45d3-89bc-0026455624cc · outbound

This paper cites Then express a, b, c in terms of these values and use linear programming or symmetry arguments to maximize |a| + |b| + |c|.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Then express a, b, c in terms of these values and use linear programming or symmetry arguments to maximize |a| + |b| + |c|

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:8334716e0c4652d1f56ec55f7a581980245af38ff9a100c4599c0db81d21316c

Observation 7c50f582-1be0-47d3-8b2c-b8affaacb8a1 · outbound

This paper cites Apply the method of Lagrange multipliers to maximize the linear functional |a| + |b| + |c| subject to the quadratic constraint.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Apply the method of Lagrange multipliers to maximize the linear functional |a| + |b| + |c| subject to the quadratic constraint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:f9d84878e423b7304f70f610b140ab83a2b82aa472b4d05bb0acf2adb54732a9

Observation 64a72dcb-d871-4c8d-942c-993140c365fe · outbound

This paper cites However, this may miss the global maximum if the optimal polynomial is not symmetric or has non-zero a and b.

Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this may miss the global maximum if the optimal polynomial is not symmetric or has non-zero a and b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:922e2b2894e9176ea6974ed6ce52e9f919addeff6fe58fa6c0c933945e344fa2

Observation 2bd6b323-e61e-4f1b-aad7-04ed6e6a14da · outbound

This paper cites Scale and shift the Chebyshev polynomial to satisfy the bound|P(x)| ≤1and compute the coefficients to find the maximum of |a| + |b| + |c|.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Scale and shift the Chebyshev polynomial to satisfy the bound|P(x)| ≤1and compute the coefficients to find the maximum of |a| + |b| + |c|

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:18b2f907185f33f69cbd99916483764bb1debe3d46919e378c77e0f5549753ed

Observation 06aea44a-95a3-4876-8b8e-324c355a9784 · outbound

This paper cites propose-select-think.

Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:7abf496973adb3f7f36e6b6b3fbe18e0607d0647e12fe3d28ca3de561ac62d86

Observation db848297-b0a0-4389-a502-48367449eb9e · outbound

This paper cites - PointPis 4 units away from the circle, so the distance fromPto the centerOis6 + 4 = 10.

Hint-Guided Diversified Policy Optimization for LLM Reasoning - PointPis 4 units away from the circle, so the distance fromPto the centerOis6 + 4 = 10

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:48895aa0cedde80f2f282c72ebcd1038a185048d8fb600cf3031de8061807d99

Observation ea92010a-ee05-4c91-b1eb-61b6c1741a0d · outbound

This paper cites But sincePis 4 units away from the circle and AB is parallel to ← →OP, the perpendicular distance fromOto ABmust be 4 (as 8 would place AB outside the circle).

Hint-Guided Diversified Policy Optimization for LLM Reasoning But sincePis 4 units away from the circle and AB is parallel to ← →OP, the perpendicular distance fromOto ABmust be 4 (as 8 would place AB outside the circle)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:61778ecde4536587adbf9ae1270adde24f355525551d9ef653c5e95f317c817f

Observation a49c998f-77bf-426b-8104-e315f03ea951 · outbound

This paper cites - The chord ABis parallel to the x-axis and 2 units below the x-axis (since the distance from Oto ABis 4).

Hint-Guided Diversified Policy Optimization for LLM Reasoning - The chord ABis parallel to the x-axis and 2 units below the x-axis (since the distance from Oto ABis 4)

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:13b8d1a365bbfcb0a6481fc3fd576263099c44fa30272cd184a885ead9b6a78a

Observation 12e5b615-a66f-47a8-95a3-63c9337fd251 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:b847fed5542ebd57e6096943e7d9a2a6882b159342f5b464a84e986ceda0998d

Observation fb498f7e-d7a3-41f3-bc09-7119a2cc1a85 · outbound

This paper cites 19 Case 2 (Generation Model: Qwen2.5-Math-7B-HDPO) Question: CircleOhas radius 6.

Hint-Guided Diversified Policy Optimization for LLM Reasoning 19 Case 2 (Generation Model: Qwen2.5-Math-7B-HDPO) Question: CircleOhas radius 6

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:c883fb8f88599ae443cc4edacf76dab945043334dc35dbfcf9320b388817cfff

Observation 5a4b1468-b2e5-428d-add9-19958630381e · outbound

This paper cites SinceABis parallel toOP and the distance between them is 2, the perpendicular distance fromOtoABis either4 + 2 = 6 or4−2 = 2.

Hint-Guided Diversified Policy Optimization for LLM Reasoning SinceABis parallel toOP and the distance between them is 2, the perpendicular distance fromOtoABis either4 + 2 = 6 or4−2 = 2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:986c2c306c168c8c28e3e11aa029ada108ef8c26d495160cd244d7f9fe65fbf5

Observation a1a27ee0-65a8-486e-8b41-3e48c3cc7f62 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:32875221fe7c7523b07291670414533550c7be0ad3492e67c0460182ce1aea8c

Observation 7e1bcde0-f444-4255-906d-6ae29d70119e · outbound

This paper cites Since chordABis parallel to ← →OP, it is horizontal, and the distance betweenAB and ← →OPis 2, soABis either aty= 2ory=−2.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Since chordABis parallel to ← →OP, it is horizontal, and the distance betweenAB and ← →OPis 2, soABis either aty= 2ory=−2

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:13193f8c4aeee54f55d7bdf17c8686aff1503d79aebd7f6a9e864d7bff8fa5a5

Observation d04f6e3f-c5f9-456a-b7c2-60bc57b60bfa · outbound

This paper cites propose-select-think.

Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:3f1b1993e8b33e4c343234c0858c3a778af2d6d969fc2162cda761c5c1c6e196

Observation 252f6b2d-5468-478d-8941-fa2f24bf7eb3 · outbound

This paper cites Then apply the sum of cosine series formula for angles in arithmetic sequence, simplifying the resulting expression using symmetry and periodicity of the cosine function.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Then apply the sum of cosine series formula for angles in arithmetic sequence, simplifying the resulting expression using symmetry and periodicity of the cosine function

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:5f64a9bd687f0f590405996c4a6c929399194884761ca0c5bbd6e48f7c7ccf15

Observation c4f4b924-04b7-4722-aad8-9cd5106fbf03 · outbound

This paper cites However, this approach lacks precision and relies on approximation, making it unsuitable for exact computation.

Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this approach lacks precision and relies on approximation, making it unsuitable for exact computation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:bb108cfd9f97f4b2163fc7cd813552c33984b33ba85d3115f275d28ba1826b49

Observation aa14e551-b1c2-478d-8bb3-6488dd20bb50 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:05f9d916169f8ee7f7b7755dc3667f8815504fe2f9fbcc0849a49d271ae62169

Observation 55e0e174-2b13-4aaa-a159-b94a4b861385 · outbound

This paper cites </Candidate Solutions> <selected>[1]</selected> <thinking> We are given the sum: sin2 4◦ + sin2 8◦ + sin2 12◦ +· · ·+ sin 2 176◦ This is a sum ofsin 2 θforθ= 4k ◦ wherek= 1,2,.

Hint-Guided Diversified Policy Optimization for LLM Reasoning </Candidate Solutions> <selected>[1]</selected> <thinking> We are given the sum: sin2 4◦ + sin2 8◦ + sin2 12◦ +· · ·+ sin 2 176◦ This is a sum ofsin 2 θforθ= 4k ◦ wherek= 1,2,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T10:55:33.276019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:55:33.276019Z digest=sha256:56f90ea8a6c8c57f5f3018abecd65588e65d2cb108fb5329fa5150312be8d67b

Pith citing papers

No inbound Pith citation observations are available.