1 comments

  • qiuwu 2 hours ago

    It seems benchmark still far lagging beyond real abilities of LLM.

    Does that count? Arena is real people, but only common parts, or might be a standard of another scope, i.e. personal interest or tast.

    If comes to scientific side, the measure is really easy strap into a bias.