Pith. sign in

Paper Citation Record · LEDGER

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

As of 22 July 2026, this Paper Citation Record lists 31 of 31 outbound references and 16 inbound Pith citation observations for arXiv:2510.18471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18471 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T05:16:28.008746Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-20T06:30:07.809122+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T14:30:20.959431Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact24
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb50c0e3-7c6a-4401-9a99-147dea01ef50 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.609566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:ca406190fa2948e54f49e105898c9f17e3ff6795c3e0291e6a16e5a52e65c952

Observation 133023d3-99be-4eae-9140-1fb305078c90 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.602730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:a0f1d4f4d0adb96fdf4c882ba3ff56f13aa48135ecdd0b9cc1f1809dc8990e3f

Observation 35c381a8-345a-4eea-b64f-bc280f5d75ff · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.624448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:eb069a768b2493c778107d6f7212b887baed93759875269324c8efedfc6ffb1b

Observation 9f1ccf0c-4e9f-499f-85d0-45291f2e4ac4 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Process Reinforcement through Implicit Rewards

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.579753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:993db4073dcec03e74cec28effc518941cf46b5374164dd7a76de5c0d3d94990

Observation 19022fa2-d24f-4201-97be-b9d8355ac07a · outbound

This paper cites CodeScore: Evaluating Code Generation by Learning Code Execution.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeScore: Evaluating Code Generation by Learning Code Execution

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:54.620410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:4f849c7d6894b6c7fe2d484fcd0e59ec4336413f8fb41c6a6aab395899e36a96

Observation f3ba0fca-58f2-4a83-98fe-1e1ceb21b23b · outbound

This paper cites A Survey on Code Generation with LLM-based Agents.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment A Survey on Code Generation with LLM-based Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:02:21.366092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:3d289f702dc84777806a2418f8f21eab28dc15820713019e570159666c1411b3

Observation 33cd7339-6bfe-4775-ac0e-7e589667d67e · outbound

This paper cites ReCode: Reinforcing Code Generation with Reasoning-Process Rewards.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T05:20:54.545492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:fcc61b0031f8831698a5fbab803cd487a2f553981a475a48feade68e66df6f7d

Observation a92c6490-3bd4-4fd2-9249-c8876bf892ad · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment MiniLLM: On-Policy Distillation of Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.589881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:c32e92bff82ea27fe6c748723cd951e4b45c227b6fcff12a34966d4c929640c7

Observation 42fdce16-fd98-40e5-8748-dc2aaa83dbe6 · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment The False Promise of Imitating Proprietary LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:54:31.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:b5a1706795555a6f3da055fd3f5fec1733304ece4308076176447795df027f5c

Observation 86d251eb-e19e-42fc-aca2-e7b41d675a63 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Teaching Large Language Models to Reason with Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.606098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:fd2ebc29c8fa165e529dcaaa3e7987d425a463bee40681c12730b5b903f8532e

Observation 4509fc09-955f-49be-9835-cd4b484f275c · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Skywork Open Reasoner 1 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.576707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:e19df6bfa5a998674dc0883bdeba94c721ef8b561bd4b2922c5ca09832f6c6cf

Observation c92de9ca-b71f-443b-a17b-025fe33c0ba0 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Measuring Coding Challenge Competence With APPS

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.613065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:363a889a9080a397232ca959311baf121c9397a2ae5a5de160f78dc395a35736

Observation bc0ebe0d-b6de-4509-8776-3e848d148cc1 · outbound

This paper cites Designing and interpreting probes with control tasks.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Designing and interpreting probes with control tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.141010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:9b3cd1b2cf15cc44cf55bbab175da281b41967abdfa28e38d42b3b130f5a099f

Observation feda402c-46a4-464d-9774-51e523e764a1 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Qwen2.5-Coder Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.522277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:45a7dc9a6f3210bbc0f8d75915df8932d05f95230609781d1677fb9ed5337ac1

Observation 428da1f5-5f6a-42a0-9113-5b1de683fc1d · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.599756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:61d7274b5b9acc333e0ef0bdc5f54ba2e809ddc4e3f17c09d17da3f95f1ecc1f

Observation fae20ebe-62cb-41d4-8060-b9a257df0488 · outbound

This paper cites SEED: customize large language models with sample- efficient adaptation for code generation.CoRR, abs/2403.00046, 2024a.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment SEED: customize large language models with sample- efficient adaptation for code generation.CoRR, abs/2403.00046, 2024a

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.555648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:a0fe0a696e3da28a086c2d215235965ad4fef80161fa0939f54adb7b689e3176

Observation 883330c4-a3b1-4746-b00f-3c2b18ea02a0 · outbound

This paper cites How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.586746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:8105681164124feb266d1d53eca32e12d723b7119ec3927f4db92057cd7e402b

Observation 498c88f4-de7e-41d7-b0fb-2772d6d3bf35 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment TACO: Topics in Algorithmic COde generation dataset

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.527865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:e58a59bdf43c0865ef7b12e3614294b39921d1707d2183ace3f3745a341b9bc4

Observation 67f68427-9c13-4e8f-8a61-4f441c4caee8 · outbound

This paper cites GPT-4 Technical Report.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment GPT-4 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.616387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:6def00627ff0a7969bfe46dccf4068208ed76780e45b3b6011bc4a988f1f7d2b

Observation d75a2d4a-d82e-47e6-924a-1b12a506e19b · outbound

This paper cites Code Llama: Open Foundation Models for Code.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Code Llama: Open Foundation Models for Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.631628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:2ec4ef26de6c0f30877785394cfe08118a79b53dc0641ada0e5d67815b0080fc

Observation f338b61f-d491-4256-adb8-294aac569a5d · outbound

This paper cites Proximal Policy Optimization Algorithms.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Proximal Policy Optimization Algorithms

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.627938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:5b8b80fd529c5de0cd85e6675c6c5a56c68cd95df2b2794a8a011fa3d9059d4d

Observation 1ef83e66-5b18-4934-a089-3cab40d05de3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.559062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:fe2a92e9372cfaecdcd89fdf12721f5bd82a288784dad8513b91ed8bceea85fe

Observation 7b23b830-ed05-4bf7-99dc-b4222ef13cee · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment HybridFlow: A Flexible and Efficient RLHF Framework

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.549662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:019e3e8a58e2e2d2a5014e048f4047524e767281baaa5090bd0af212e8cedf98

Observation bf918166-6fc3-4071-8ec0-d28d694f7339 · outbound

This paper cites CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.583092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:4cf030a30b38d4ab39fc908bbbfd6903a69969b6795b8a69785aa7cce5048e6b

Observation 7bd90871-6fa5-4941-b7e0-192da2733377 · outbound

This paper cites CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.573655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:1525cd1c5c227a743049c00bd25f8b9173c41a63ed49d7748eba8e0c3f0747af

Observation c8a3f9db-b5ba-4130-b6c0-afbef25450ab · outbound

This paper cites Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Co-evolving llm coder and unit tester via reinforcement learning.arXiv preprint arXiv:2506.03136

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:54.562712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:33b776001d4dd3ae3ab143e89275dde538aae501ad37a1e0386fc33d75b7a8e3

Observation c5509687-8d24-4de0-8b8a-917d97791d98 · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:20:54.593415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:80a4e0dce62f29c671c511c190a288249cbdeec7f5a034ee56133e05fbdda459

Observation d6eb99f1-bcfd-41d2-bf6a-f13127b68e17 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment A Survey on Knowledge Distillation of Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:20:54.566348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:7b7ddda6773190d9786e08b66f8c1e0f5fcf9062bd9b04775d8e6b3d9d81b461

Observation 835bd351-177b-4e26-abc7-1012ea600259 · outbound

This paper cites Here's a step-by-step approach to achieve this:1.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Here's a step-by-step approach to achieve this:1

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.133531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:1e7d4af67541b51a864b56d8c9e49935c717f2e21d7c6309ed4fce520cfde8f7

Observation fb4308cd-a36e-4705-927c-258161822413 · outbound

This paper cites Initialize a Dictionary to Track Blocks: Use a dictionary to map the top-left corner of each 2x2 block to the count of black cells in that block.2.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Initialize a Dictionary to Track Blocks: Use a dictionary to map the top-left corner of each 2x2 block to the count of black cells in that block.2

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:25:55.136519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:ed23fc5d889589853adfe29633d751d6a991ba6e3207e2aec2f127fcabd36918

Observation 815fb496-5bf3-4a86-b4e5-5b659f91ef37 · outbound

This paper cites an unresolved cited work.

CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-18T05:25:55.138529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-18T05:16:28.008746Z digest=sha256:90e33c254938721841937fb0d26aa66cad60e30fd03281a909d1f7c609bbf732

Pith citing papers

Observation d84a2454-7244-45eb-b109-e7f1f8a552cd · inbound

Think Anywhere in Code Generation cites this paper.

Think Anywhere in Code Generation CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:18:25.612924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T23:16:42.782431Z digest=sha256:c7914b1d4752e391e82da9ba4beae186d708420ca8a1ef3fa83262b72888eda9

Observation 6a76c5aa-6361-4f47-8d5e-46524980184e · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:38:18.780983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:8a4c111a71e11725b130793e7e50a4fb6b7cad0bca46672d3af1898c6bb70764

Observation 830188b8-49f9-4487-9787-53a440b74aeb · inbound

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy cites this paper.

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:53:11.753782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-13T19:51:08.304741Z digest=sha256:d8cef8db572355694f1049c835a6d21b1e332a6b6318d0f76385e30fa22b19d8

Observation b9f4dcbc-5880-46cb-8691-fba541979f9c · inbound

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models cites this paper.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.478225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:672633f27b6a4c19eef12f68cfb5c484e9bd30951b2096d2a8c839d25a3a3de7

Observation dfa64b51-5155-4abd-b316-84f35d5c18d6 · inbound

Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs cites this paper.

Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:10.469171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-08T09:11:20.733960Z digest=sha256:2cba72547e247704896420cd465dea5d5e261973697002c8c21932d305b2b10b

Observation bb4b95a2-b255-4026-908e-42ee84ae03c4 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:43.108008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:a8db2f1dde4020108a500617a7a870ef82bc0245e2e9a2b4ddfcc909432e8ea2

Observation 7c6a563f-8d8c-4ba8-ad5f-9cff54d5c9dc · inbound

Code as Agent Harness cites this paper.

Code as Agent Harness CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:14.219194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-05-20T10:54:54.558241Z digest=sha256:f712d30a1ddb6b17ecf45e15f072e5c917cb856c076aa9e3b83c7d4a66a2221c

Observation a33f4fe7-b7ed-4d90-9ab7-71c8191ca18e · inbound

Distilling Game Code World Model Generation into Lightweight Large Language Models cites this paper.

Distilling Game Code World Model Generation into Lightweight Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:04:44.660858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-30T13:58:37.956333Z digest=sha256:00a58322ccfdbca203e507c1f821a531b57c7fd00b7175761d396d21ad2ab2c3

Observation 737d4a5e-111d-4cbb-b10a-0c5d1446197f · inbound

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback cites this paper.

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T14:53:32.279726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-29T06:11:06.960389Z digest=sha256:f238ae29cfe7bd906f5972aaafbf4332b000a138f56dc0f4699dfa9a90f330d9

Observation 45b56bcc-7699-4426-ac98-e248ca395c1b · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.774024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:1dca55a739c15182a137ecc406271981d1628d2510881ed9cfb70a72f4d63dc9

Observation aa580c9a-1676-4f23-8d9d-52405be4a5f2 · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.933247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:2f75b9321d2f8570b1adaaf4cce58b7e2118681e9d296131af2db2e2af17387c

Observation 54136184-f7a3-452a-b109-c60542fdb2ef · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-27T13:10:55.917848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:a2fba5a25c1a58e6d540384867bf4999178a3bb9b25033feb745a597ccce754b

Observation 1fa12586-08a0-4488-af79-31c964d593d0 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T12:58:08.768615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:4b29d3d84629d4f8d25503e6c49904d061d678d1aa856863b8b300ef867ec62e

Observation e2839e89-9d81-4a5b-9f33-b3947090af37 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:00:19.806171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:a7e9888fbde5224d62aa12dea805e51035beb6c66d9cc1939a836cca96b31a0d

Observation a9389d13-50cf-41b5-ad78-a794dfa3a7fa · inbound

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study cites this paper.

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T04:19:33.972007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-20T06:30:07.809122+00:00.

source=pdf_text observed=2026-06-26T17:07:21.486960Z digest=sha256:db5b9617be4a2b2601d9b46928cc7b252c78dc2e67198f6e8580c209463780bd

Observation 401e2898-b4b6-464c-8117-91fcc19a8df3 · inbound

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment cites this paper.

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T14:30:20.959431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:30:20.959431Z digest=sha256:a23cb4e0479907d43a26527a6182f74a32a601d0baf30bd28b872294be8247a3