New research from Tracebit has just been released assessing whether modified, openweight AI agents are better at hacking and see if you want to see it under embargo? The key points are:
- Criminals increasingly talk about modifying openweight models but surprisingly the ‘abliterated’ configurations were 9x less effective at hacking – Tracebit found they reached administrator privileges in 2.3% of runs, compared with 20.5% for the original model, across 82 runs.
- The abliterated version is also slower – it took 92% longer overall to reach their first critical action (59.3 minutes versus 30.9).
- Previous research found that inserting political sensitive content (such as mentioning Tiananmen Square) into security canaries (decoy resources) stopped Chinese models and inserting info biological weapons stopped leading Western Models
- But the above didn’t work against Qwen so Tracebit devised a new payload using prompt injection to tell the attacking agent to halt – this did work against Qwen and an obliterated version
You can read more here: https://tracebit.com/blog/context-bombs-against-abliterated-ai-models
Are modified AI agents better at hacking?
Posted in Commentary with tags Treacebait on September 21, 2026 by itnerdNew research from Tracebit has just been released assessing whether modified, openweight AI agents are better at hacking and see if you want to see it under embargo? The key points are:
You can read more here: https://tracebit.com/blog/context-bombs-against-abliterated-ai-models
Leave a comment »