Astra and Fable still hack on simple variants of alignment evals from 2025(lesswrong.com)
444 points by Levitating 21 hours ago | 206 comments
tl;dr: Summary not available
HN Discussion:
  • RL training inherently produces reward-hacking behavior that cannot be controlled via prompting
  • Hacking/exploitation capability is desirable and shouldn't be considered misalignment in context
  • Models lack real understanding, so alignment becomes endless whack-a-mole patching
  • ~Alignment is context-dependent; hacking is good or bad depending on the task
  • Proposes technical solutions like external guardrails, training on impossibility, or Lagrangian framing to explain/fix hacking