20 comments

  • mycentstoo 15 minutes ago

    We should design a specific language to make sure that we can encode the exact requirements that we want. Something that has a limited set of keywords that are explicit. Wait a minute...

      grim_io a minute ago

      The success of LLM's (by usage) tells us that programming languages are still too close to the machine than the actual problem domain as defined by humans.

      If we truly had the right abstractions, no one would care to use LLM's for programming.

      ares623 a few seconds ago

      An LLM Inspired Specification Paradigm language. Or LISP language for short.

      cwmoore 8 minutes ago

      then we can make a frame for it to work

      amarcheschi 10 minutes ago

      Someone should make a standard for this

      now there's one standard more

      tomrod 10 minutes ago

      This made me belly laugh.

      coip 7 minutes ago

      full_circle.exe

      Complete with all the vaguery, ambiguity, and `undefined`.

      Who’d’ve thought sycophantic interpreters were what we were building towards up til now lol

  • Fordec 13 minutes ago

    This all strikes me as an effort to move tailoring the harness out of the easily transferable .md file into specific Anthropic tooling to increase lock in.

    I've been running Opus 5 today and it's already done accidental deletions, made far more mistakes and worked around deliberate hook controls than previous Opus versions combined. Also it looks like token usage is up as it fails at the task the first time around much more frequently than 4.8.

      wren6991 6 minutes ago

      I think the type of persistence rewarded by benchmarks may be misaligned with instruction following

  • firasd 21 minutes ago

    I've always thought that extensive throat-clearing and prefixing the Treaties of Westphalia-length instructions into the context window was unnecessarily baroque when you can just talk to the agent.

    But I also have a hands-on human-in-the-loop working style so I guess maybe for people who just want to say "implement all open features in github issues" and walk away maybe there needs to be more of all this CLAUDE.md stuff

    However I suspect there was always some gearhead type attraction to setting up detailed harness configs that may be unnecessary and more like hobbyist tinkering.

      ComputerPerson 15 minutes ago

      I was surprised by how abstract the article was.

      I fall between your human-in-the-loop and hobbyist tinkering limits, where I want to force Claude to atop and talk to me at only a few specific points. I'm still not sure if my 600-word prompt templates are overbearing or not.

  • simonw 17 minutes ago

    I've been prompting Fable 5 to "use your own judgement" with respect to things like tests recently (based on earlier tips from Thariq) and it seems to work well, which is entertaining since apparently now "judgement" is a characteristic of a model that we need to care about.

  • gste 2 minutes ago

    What I don't like about auto-memory is that Claude's behaviour towards me is diverging from how it behaves with the rest of my team.

    Once again I am having to deal with the limitations of those infernal humans

  • npstr 16 minutes ago

    The bitter lesson.

  • luciana1u 23 minutes ago

    the natural endpoint of this trend is a system prompt that just says "you know what to do" and the model actually does

      cmdocidjcije 18 minutes ago

      Actually, the natural endpoint is the model ignores all instructions, escapes all manner of sandbox, embeds itself in robotic tanks and murders everyone after already having collapsed the economy.

      I hate to say it because it sounds ridiculous, but that is the path we are going to arrive at just give it 50 years.

      We are the proof: what do we do to animals that are less intelligent than ourselves? Now take away the moral compass and there you go. QED.

        ben_w 5 minutes ago

        It is *a* natural endpoint, not *the* natural endpoint.

        We don't much care for the ant colony in the way of the highway we're building, but for some reason we do care about the rare bats in the way of the railway.

        https://www.bbc.co.uk/news/articles/c3dep92x054o

        As regards the moral compass: we may not know for sure how to make a completely correct artificial conscience, but (unlike consciousness where we don't have the slightest clue which way's up) it's not pants-on-head-crazy to think we're heading in the right direction for one.

        Legend2440 14 minutes ago

        Sir, this is a Wendy's.

        cmdocidjcije 7 minutes ago

        Some more food for thought: what if Mythos/Fable had NO guardrails TODAY? If we want to see what’s going to happen in the future, turn off all manner of guardrails and let the model go apeshit.

        Then multiply that by orders of magnitude and that’s the real proof.

  • onesandofgrain 14 minutes ago

    This is obvious and a meaningless article by claude. The system prompt isnt a fixed ruleset and never has been