The way he proposed to deal with "How urgent is this customer support message?" is "a score over ordered levels".
I'm not sure that's a good way to use Jev for this use case. If I had this problem, I would ask Jev to answer several different yes/no questions about each message, and then use the probabilities as inputs into a logistic regression model that predicts urgency.
If you ask the model specific yes/no questions which can be answered reasonably objectively from the input, I think the answers are going to be more stable over successive generations of models.
e.g. if you ask 'Is the customer angry?' I'd expect that answer to have high agreement between models and between models and humans. But directly answering the 'is it urgent' question is much harder. (Although I suppose you can try to put the rules in the prompt.)
Calibration has to be measured on your dataset. Just because the probs sum to 1, does not make Jev or Jev-like models claibrated. For folks interested in digging deeper into calibration, studying ad click prediction models (where calibration is super important) is a good place to start.
The way he proposed to deal with "How urgent is this customer support message?" is "a score over ordered levels".
I'm not sure that's a good way to use Jev for this use case. If I had this problem, I would ask Jev to answer several different yes/no questions about each message, and then use the probabilities as inputs into a logistic regression model that predicts urgency.
If you ask the model specific yes/no questions which can be answered reasonably objectively from the input, I think the answers are going to be more stable over successive generations of models.
e.g. if you ask 'Is the customer angry?' I'd expect that answer to have high agreement between models and between models and humans. But directly answering the 'is it urgent' question is much harder. (Although I suppose you can try to put the rules in the prompt.)
Calibration has to be measured on your dataset. Just because the probs sum to 1, does not make Jev or Jev-like models claibrated. For folks interested in digging deeper into calibration, studying ad click prediction models (where calibration is super important) is a good place to start.
Spot on. A slightly less accurate but calibrated model inspires far more trust than a 'perfect' one that's consistently overconfident.
Humans don’t use “quietly” in normal parlance. Come on man, try harder.
Edit: Dude, you didn’t even use the thing you’re talking about? It’s on OpenRouter. Do better!
Edit 2: OP is a ~60 day old account, only other (positive) commenter is a ~48 day old account. Sus.