7 comments

  • jrflo 20 minutes ago

    As an overall pro-AI person... Please don't let it do the entire UI design for your site. 99% of this site is completely useless.

  • mkagenius 44 minutes ago

    > What the two layers say

    > Stated findings

    > Findings derived from two curated layers: which model cards mention each benchmark, and which scores could be read verbatim from those documents. Each finding names the evidence behind it.

    This is 100% AI generated but the problem is it's difficult to understand - what layers is it talking about, is it the llm model layers or what.

      wolttam 31 minutes ago

      I’m getting good at reading LLM-speak (which is probably sad)

      It means two “layers” of information access/validation.

      Sounds like it first made a little map of model <-> benchmark, then went and filled in the score boxes.

      Definitely not LLM layers

  • claiir 23 minutes ago

    > This measures vendor attention, not benchmark quality

    This text on this page is so aggressively LLM-written (Claude) I am struggling to understand what I am even looking at.

  • ttul 41 minutes ago

    I am waiting with bated breath to read, “load-bearing” somewhere… The latest models are very capable, but sometimes they seem to get so deep in the details that they lose the overall plot.

    What the hell is the point of this page? Can you put in a single bit of human prose explaining why it exists and what we are supposed to learn?

  • augment_me 26 minutes ago

    Good idea, would be interesting to cross-examine the benchmarks, but the page information is completely obscured by the AI slop. The benchmarks comparison and should start immediately instead of having random completely arbitrary complex headers and labels

  • bilbo-b-baggins 40 minutes ago

    Broken on mobile safari