Kaan · Article
2025-11-17
Penalties Beat Rewards: RL Lessons from Cursor Power Users
Negative feedback outperformed positive reinforcement. Reward design beat algorithmic complexity. Richer observations led to better generalization.

Penalties > rewards. Reward design > algorithms. Richer observations > edge cases.
What happens when you put Cursor power users in a room to build RL environments? I organized a session last Friday that surfaced three clear learnings:
-
Negative feedback outperformed positive reinforcement. Teams using strong penalties for bad actions saw noticeably faster improvement.
-
We found reward design beat algorithmic complexity. Domain knowledge in the reward function mattered more than the algorithm itself.
-
Richer observation spaces led to better generalization. Agents trained on varied inputs handled edge cases without extra work.
Thank you to Cursor for sponsoring, and to everyone who participated. The debugging conversations were just as valuable as the pizza.
Looking forward to the next session.
*On what topic would you want to prototype with the group? Drop your ideas below or DM me.