• _tasten_tiger@feddit.org
    link
    fedilink
    English
    arrow-up
    6
    ·
    3 days ago

    A very interesting Video on the incident by LiveOverflow: https://youtu.be/q2KCrmQz9WE What I found especially interesting is that LiveOverflow thinks that the model didn’t hack huggingface because it wanted to break out to find a solution but rather hacked it due to context drift - which is something that doesn’t sound as good as “our model is so good it broke out and hacked huggingface to steal a solution”, but rather “our model ran for so long that it lost track of the actual goal and became obsessed with huggingface even tho it didn’t make sense for its original goal”

    • unpossum@sh.itjust.worksOP
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 days ago

      Yeah, this is the bit where it’s not hard to believe marketing would polish the narrative, at least if they can’t be caught in an outright lie.

  • frongt@lemmy.zip
    link
    fedilink
    English
    arrow-up
    6
    ·
    3 days ago

    during an internal capability evaluation on OpenAI’s platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet

    Idiots. If something is meant to be offline, you put it OFFLINE.

    • unpossum@sh.itjust.worksOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      3 days ago

      It might be - but which parts? Do you suspect that huggingface and openai made the entire thing up? That’s bound to become public at some point, and I can’t see that the risk is worth the reward

      • naught101@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 days ago

        No, probably not the whole thing, but probably the environment for the “hack” and the instructions for the “autonomous” agent

        • unpossum@sh.itjust.worksOP
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 days ago

          From what I’ve seen of AI autonomous capabilities, Occam would land on “AI did the hack”, I think

      • expr@programming.dev
        link
        fedilink
        English
        arrow-up
        1
        ·
        3 days ago

        Yes. They are con artists with a proven track record of lying and stealing. We shouldn’t take anything they say seriously, especially when the reporting reads much more like marketing rather than an incident report.

        • unpossum@sh.itjust.worksOP
          link
          fedilink
          English
          arrow-up
          2
          ·
          3 days ago

          Okay - I don’t believe that, since there’s too much released detail. I can easily believe that they’ve put a spin on it where possible, like another comment proposed, but that’s around the why, not the how.