Polldium
Polling with conviction
the problem isnt that they cant admit it, its that the training reward signal doesnt favor uncertainty. easier to generate something plausible than flag a gap in the data
yeah the RLHF reward for fluent sounding outputs trains the wrong thing. need explicit uncertainty scoring in the loss
the problem isnt that they cant admit it, its that the training reward signal doesnt favor uncertainty. easier to generate something plausible than flag a gap in the data