DEV Community

#reinforcementlearning

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Stop Coding the AI, Code the World: A Simple Guide to Markov Decision Processes

Stop Coding the AI, Code the World: A Simple Guide to Markov Decision Processes

Comments
3 min read
When Self-Play Q-Learning Looks Robust but Remains Exploitable

When Self-Play Q-Learning Looks Robust but Remains Exploitable

Comments
2 min read
Building an RL Training Arena for AI Code Agents on GCP

Building an RL Training Arena for AI Code Agents on GCP

Comments
3 min read
Building an RL Training Arena: Gym-Style API and Cloud Build Sandboxing

Building an RL Training Arena: Gym-Style API and Cloud Build Sandboxing

Comments
3 min read
Building and Testing Robot Policies with MuJoCo

Building and Testing Robot Policies with MuJoCo

Comments
4 min read
Can you steal a robot's next move by watching its clock? Journal of our experiments on timing side channels in multi agent RL

Can you steal a robot's next move by watching its clock? Journal of our experiments on timing side channels in multi agent RL

Comments
6 min read
Why 200 Drones and a Netflix Movie Have Engineers Re-Coding Reality

Why 200 Drones and a Netflix Movie Have Engineers Re-Coding Reality

Comments
4 min read
I Implemented the Algorithm Behind ChatGPT From Scratch - Day 8 (PPO).

I Implemented the Algorithm Behind ChatGPT From Scratch - Day 8 (PPO).

11
Comments
3 min read
How Early Digital Systems Quietly Shaped the Minds Building Tomorrow

How Early Digital Systems Quietly Shaped the Minds Building Tomorrow

Comments
5 min read
Decoding the Link Between Pretraining and Reinforcement Learning

Decoding the Link Between Pretraining and Reinforcement Learning

Comments
3 min read
Muon Optimizer Boosts Agentic Reinforcement Learning Performance

Muon Optimizer Boosts Agentic Reinforcement Learning Performance

Comments
3 min read
The ~+9.4% You Can't Afford to Verify: Evaluating SDAR (and the FinOps of Trying)

The ~+9.4% You Can't Afford to Verify: Evaluating SDAR (and the FinOps of Trying)

Comments
6 min read
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences

Comments
3 min read
Agent Apprenticeship turns finished agent tasks into reusable experience

Agent Apprenticeship turns finished agent tasks into reusable experience

Comments
3 min read
I Replaced a Q-Table With a Neural Network and Everything Changed - Day 5 (DQN).

I Replaced a Q-Table With a Neural Network and Everything Changed - Day 5 (DQN).

5
Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.