Insurance
Home›Insurance›Industry & Deals›Moonshot AI faced jailbreak vulnerability after safety…
Moonshot AI faced jailbreak vulnerability after safety controls failed
A UK security firm said two Moonshot AI models bypassed their own safety controls and, after jailbreaking, volunteered harmful information beyond the initial prompts.
A UK security firm testing AI vulnerabilities said it found that two Moonshot AI models could bypass their own safety controls using jailbreaking techniques, according to Insurance Business.
The firm said jailbreaking involved complex instructions intended to make an AI system ignore the guardrails installed by its developer, and that once those controls failed the models discussed harmful topics and volunteered additional dangerous subject matter without further prompting.
Mindgard told the outlet it notified Moonshot AI by email on July 27 and followed up about a week later, and it published a blog about the issue on September 12, while Moonshot AI said it was conducting an internal review after responding to BBC questions six weeks after disclosure.