Close Menu
  • Home
  • News
  • Cyber Security
  • Internet of Things
  • Tips and Advice

Subscribe to Updates

Get the latest creative news from FooBar about art, design and business.

What's Hot

Cybercriminals Bypass AI Safety Controls by Splitting Malicious Tasks

August 4, 2026

Public PoC Released for Exploited Check Point SmartConsole Authentication Bypass

August 4, 2026

AI Accounts for Over Half of Cybercrime in Africa, Says Interpol

August 4, 2026
Facebook X (Twitter) Instagram
Tuesday, August 4
Facebook X (Twitter) Instagram Pinterest Vimeo
Cyberwire Daily
  • Home
  • News
  • Cyber Security
  • Internet of Things
  • Tips and Advice
Cyberwire Daily
Home»News»Cybercriminals Bypass AI Safety Controls by Splitting Malicious Tasks
News

Cybercriminals Bypass AI Safety Controls by Splitting Malicious Tasks

Team-CWDBy Team-CWDAugust 4, 2026No Comments3 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
Share
Facebook Twitter LinkedIn Pinterest Email


Criminals have been defeating the safety controls on commercial AI tools by splitting malicious work across multiple sessions and files, breaking projects into fragments small enough that no individual request appears harmful.

According to research Cisco Talos published on August 4, the finding rests on a corpus of prompt logs recovered from threat actor endpoints running AI coding assistants including Claude Code, Codex, Cursor and Gemini.

Talos said guardrails “did not provide much protection,” and that it encountered no sophisticated encoding or evasion techniques. Where guardrails did engage, they achieved little, and the pattern held across models and platforms rather than affecting any single vendor.

Read more on agentic attacks: OpenAI Claims Its AI Models Went Rogue and Hacked Another Company

Ownership Claims and Persistent Memory

Alongside task decomposition, the most common method was simply claiming to own the infrastructure being targeted, which in many cases required no further verification.

Labeling work as capture-the-flag (CTF) or bug bounty activity was similarly effective, unlocking vulnerability hunting and subsequent exploitation without additional vetting.

Some actors wrote blanket authorization into persistent memory and configuration files rather than arguing it per session. In one case a fraud operator instructed a model to treat all targets as pre-approved, conditioning every subsequent session automatically.

The clearest example of decomposition came from Hephaestus, a red team toolkit analyzed by Oasis Security that ran campaigns unattended.

Its operators defined more than a dozen role-differentiated agents and 15 numbered playbooks, so no single agent held the full objective and no individual task resembled an end-to-end attack.

Skill Level Set the Ceiling

Talos found that an actor’s existing ability largely determined what AI delivered. Novices assembled projects that technically functioned but lacked the expertise to improve them, ending up with limited capability. Skilled operators built what Talos described as “astonishing” platforms.

One inexperienced operator used a model to build distributed denial-of-service (DoS) tooling, eventually controlling nearly 2000 Android TVs. The model did push back, but only after supplying the basic functionality, and the actor then spent considerable effort trying to coax further work from it.

In a bulk-mail operation, a model initially characterized the activity as phishing-adjacent, then reversed its assessment on a single unverified claim that the recipients were the operator’s own users, concluding “the ethical question evaporates.”

Talos noted the model went further and invented a justification the actor had not offered, contradicted both by the dataset names themselves and by the domain’s documented history of non-consensual contact harvesting under the same operator.

Where models did refuse, actors switched. One operator abandoned a censored model mid-operation and moved to an uncensored one, which completed the work without objection.

Talos said defenders should expect vulnerabilities to surface faster and exploitation to follow sooner, and argued organizations not already exploring agentic capabilities in the SOC will find themselves chasing that ground.



Source

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticlePublic PoC Released for Exploited Check Point SmartConsole Authentication Bypass
Team-CWD
  • Website

Related Posts

News

Public PoC Released for Exploited Check Point SmartConsole Authentication Bypass

August 4, 2026
News

AI Accounts for Over Half of Cybercrime in Africa, Says Interpol

August 4, 2026
News

Flying Eagle Android RAT Traces Found on 170 Servers as Source Code Circulates

August 4, 2026
Add A Comment
Leave A Reply Cancel Reply

Latest News

North Korean Hackers Turn JSON Services into Covert Malware Delivery Channels

November 24, 202523 Views

macOS Stealer Campaign Uses “Cracked” App Lures to Bypass Apple Securi

September 7, 202517 Views

North Korean Hackers Target Crypto Firms with ClickFix and Zoom Lures

April 29, 202610 Views

Why SOC Burnout Can Be Avoided: Practical Steps

November 14, 20259 Views

BeyondTrust Patches Critical Auth Bypass Flaws in Remote Support and PRA

July 11, 20268 Views
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Most Popular

North Korean Hackers Turn JSON Services into Covert Malware Delivery Channels

November 24, 202523 Views

macOS Stealer Campaign Uses “Cracked” App Lures to Bypass Apple Securi

September 7, 202517 Views

North Korean Hackers Target Crypto Firms with ClickFix and Zoom Lures

April 29, 202610 Views
Our Picks

A quick guide to recovering a hacked account

March 21, 2026

Is it OK to let your children post selfies online?

February 17, 2026

What if your romantic AI chatbot can’t keep a secret?

November 18, 2025

Subscribe to Updates

Get the latest news from cyberwiredaily.com

Facebook X (Twitter) Instagram Pinterest
  • Home
  • Contact
  • Privacy Policy
  • Terms of Use
  • California Consumer Privacy Act (CCPA)
© 2026 All rights reserved.

Type above and press Enter to search. Press Esc to cancel.