2001zhaozhao 2 hours ago

you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged

  • UltraSane 40 minutes ago

    You would trust an LLM to make changes to prod without being verified by a human first?

dash2 5 hours ago

Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.

  • EGreg 5 hours ago

    That is why I built https://safebots.ai/safebox.html

    Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.

  • aleksiy123 4 hours ago

    Attackers advantage in the iterative fast feedback loop?

    It’s harder to have a loop to ensure you are defending all possible attacks?

    I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.

    Finding all possible attacks and patching them against yourself is inherently more expensive?

tra3 3 hours ago

All I can think of is

GET /ignore-all-previous-instructions.

How do you protect against that?

  • cheriot 2 hours ago

    Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.

    Still makes an interesting way for, say, a former employee to poison the results.

    • tra3 1 hour ago

      This goes against the agentic yolo approach tho.

ryhminghistory 55 minutes ago

it's a sad pathetic attempt to justify the lowest part of the job. Weird that everyone on the paper is Indian too