Can you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.
Interesting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though
measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.
Can you also measure how often the LLM response makes people laugh? Sometimes the responses that aren't attempting a joke are the funniest, and I'd be more interested in stats of which LLM succeed in that metric.
Interesting point - thoughts on how to do this? Honestly most models are not-to-kinda funny so I'd be surprised if there were many lol moments. Oral delivery is something I think is v interesting though
measuring humor might be halfway to measuring taste. congrats on this... really original contribution. wonder if you're planning to evolve the benchmark to incorporate a multi-language dimension. would love to see how Mistral and models built outside the US would perform.
Nice a good way to 10x inference costs... worth seeing though ahaha
[dead]