Orca-Bench: How Ready Are Language Model Agents for Oncall? arxiv.org 24 points by yruzin 6 hours ago
4di 5 hours ago looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben...This doesn't work anymore. Is there a newer link? cheriot 2 hours ago Was really looking forward to that. Hope they publish.
2001zhaozhao 2 hours ago you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged UltraSane 40 minutes ago You would trust an LLM to make changes to prod without being verified by a human first?
UltraSane 40 minutes ago You would trust an LLM to make changes to prod without being verified by a human first?
dash2 5 hours ago Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them. EGreg 5 hours ago That is why I built https://safebots.ai/safebox.htmlYour strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance. aleksiy123 4 hours ago Attackers advantage in the iterative fast feedback loop?It’s harder to have a loop to ensure you are defending all possible attacks?I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.Finding all possible attacks and patching them against yourself is inherently more expensive?
EGreg 5 hours ago That is why I built https://safebots.ai/safebox.htmlYour strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.
aleksiy123 4 hours ago Attackers advantage in the iterative fast feedback loop?It’s harder to have a loop to ensure you are defending all possible attacks?I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.Finding all possible attacks and patching them against yourself is inherently more expensive?
tra3 3 hours ago All I can think of isGET /ignore-all-previous-instructions.How do you protect against that? cheriot 2 hours ago Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.Still makes an interesting way for, say, a former employee to poison the results. tra3 1 hour ago This goes against the agentic yolo approach tho.
cheriot 2 hours ago Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.Still makes an interesting way for, say, a former employee to poison the results. tra3 1 hour ago This goes against the agentic yolo approach tho.
ryhminghistory 55 minutes ago it's a sad pathetic attempt to justify the lowest part of the job. Weird that everyone on the paper is Indian too
looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben...
This doesn't work anymore. Is there a newer link?
Was really looking forward to that. Hope they publish.
you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged
You would trust an LLM to make changes to prod without being verified by a human first?
Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.
That is why I built https://safebots.ai/safebox.html
Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.
Attackers advantage in the iterative fast feedback loop?
It’s harder to have a loop to ensure you are defending all possible attacks?
I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.
Finding all possible attacks and patching them against yourself is inherently more expensive?
All I can think of is
GET /ignore-all-previous-instructions.
How do you protect against that?
Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.
Still makes an interesting way for, say, a former employee to poison the results.
This goes against the agentic yolo approach tho.
it's a sad pathetic attempt to justify the lowest part of the job. Weird that everyone on the paper is Indian too