dcre · 44 points · 9 comments · 19. März · Open original
Comments
3 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
NInithisha220125. März
One thing missing from most RL environment discussions is observability during training. Single-agent envs are hard enough to debug, but multi-agent environments are a completely different challenge, reward curves tell you almost nothing about which agent failed or why cooperation broke down.
GIgizajob21. März
“An FAQ” really sets my grammar nerves jangling. “A FAQ” isn’t great either.
Maybe “FAQ on Reinforcement…” or “FAQ about Reinforcement…”
Comments
3 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
One thing missing from most RL environment discussions is observability during training. Single-agent envs are hard enough to debug, but multi-agent environments are a completely different challenge, reward curves tell you almost nothing about which agent failed or why cooperation broke down.
“An FAQ” really sets my grammar nerves jangling. “A FAQ” isn’t great either. Maybe “FAQ on Reinforcement…” or “FAQ about Reinforcement…”
ai slop