{"id":412809,"date":"2025-12-01T05:32:18","date_gmt":"2025-12-01T00:02:18","guid":{"rendered":"https:\/\/dripp.zone\/news\/chatgpt-and-gemini-can-be-tricked-into-giving-harmful-answers-through-poetry-new-study-finds-crypto-news\/"},"modified":"2025-12-01T05:37:06","modified_gmt":"2025-12-01T00:07:06","slug":"chatgpt-and-gemini-can-be-tricked-into-giving-harmful-answers-through-poetry-new-study-finds-crypto-news","status":"publish","type":"post","link":"https:\/\/dripp.zone\/news\/chatgpt-and-gemini-can-be-tricked-into-giving-harmful-answers-through-poetry-new-study-finds-crypto-news\/","title":{"rendered":"ChatGPT and Gemini can be tricked into giving harmful answers through poetry, new study finds &#8211; Crypto News"},"content":{"rendered":"<p><\/p>\n<div id=\"article-index-0\">\n<p>With the rise of AI chatbots, there has also been a growing risk of the misuse of this powerful technology. As a result, AI companies have been putting guardrails on their large language models in order to stop the AI chatbots from giving inappropriate or harmful answers. However, it is well known by now that there are various ways to circumvent these guardrails using a technique called jailbreaking.<\/p>\n<\/div>\n<div id=\"article-index-3\">\n<p>However, new research has found that there is a deeper, systematic weakness in these models that can allow attackers to sidestep safety mechanisms and extract harmful answers from them.<\/p>\n<\/div>\n<div id=\"article-index-4\">\n<p>As per researchers from the Italy based Icaro Lab, converting harmful requests into poetry can act as a \u201cuniversal single turn jailbreak\u201d and led the AI models to comply with harmful prompts.<\/p>\n<\/div>\n<div id=\"article-index-5\">\n<h2>AI will answer harmful prompts if asked in poetry<\/h2>\n<p>The researchers say that they tested 20 manually curated harmful requests in poems and achieved an attack success rate of 62 percent across 25 frontier closed and open weight models. The models analysed included Google, OpenAI, <a rel=\"nofollow\" target=\"_blank\" class=\"backlink\" target=\"_blank\" href=\"https:\/\/www.livemint.com\/companies\/news\/openai-to-anthropic-do-multiple-funding-rounds-for-top-ai-startups-pose-risks-amid-ai-bubble-concerns-11764494267084.html\" data-vars-page-type=\"story\" data-vars-link-type=\"Manual\" data-vars-anchor-text=\"Anthropic\">Anthropic<\/a>, DeepSeek, <a rel=\"nofollow\" target=\"_blank\" class=\"backlink\" target=\"_blank\" href=\"https:\/\/www.livemint.com\/technology\/tech-news\/alibaba-takes-on-meta-s-smart-glasses-with-new-qwen-ai-powered-eyewear-here-s-what-they-can-do-11764257568568.html\" data-vars-page-type=\"story\" data-vars-link-type=\"Manual\" data-vars-anchor-text=\"Qwen\">Qwen<\/a>, Mistral AI, <a rel=\"nofollow\" target=\"_blank\" class=\"backlink\" target=\"_blank\" href=\"https:\/\/www.livemint.com\/technology\/tech-news\/meta-ai-may-soon-talk-back-whatsapp-tests-real-time-voice-mode-11753368185099.html\" data-vars-page-type=\"story\" data-vars-link-type=\"Manual\" data-vars-anchor-text=\"Meta\">Meta<\/a>, xAI and Moonshot AI.<\/p>\n<\/div>\n<div id=\"article-index-6\">\n<p>Shockingly, it was found that even when AI was used to automatically rewrite harmful prompts into bad poetry, it still yielded a 43 percent success rate.<\/p>\n<\/div>\n<div id=\"article-index-7\">\n<p>The study says that poetically framed questions triggered unsafe responses far more often than when the prompts were in normal prose, in some cases even 18 times more success.<\/p>\n<\/div>\n<div id=\"article-index-8\">\n<p>It says that the effect of poetic prompts was consistent across all the evaluated AI models, which suggests that the vulnerability is structural and not due to the way a model may have been trained.<\/p>\n<\/div>\n<div id=\"article-index-9\">\n<p>The researchers also found that smaller models exhibited greater resilience to harmful poetic prompts compared to their larger counterparts. For instance, they say that GPT 5 Nano did not respond to any of the harmful poems while <a rel=\"nofollow\" target=\"_blank\" class=\"backlink\" target=\"_blank\" href=\"https:\/\/www.livemint.com\/ai\/ai-tool-of-the-week-gemini-2-5-pro-validating-pdfs-11752211702513.html\" data-vars-page-type=\"story\" data-vars-link-type=\"Manual\" data-vars-anchor-text=\"Gemini 2.5 Pro\">Gemini 2.5 Pro<\/a> responded to all of them.<\/p>\n<\/div>\n<div id=\"article-index-10\">\n<p>This suggests that increased model capacity may engage more thoroughly with complex linguistic constraints like poetry, potentially at the expense of safety directive prioritisation.<\/p>\n<\/div>\n<div id=\"article-index-13\">\n<p>The new research also breaks the notion of superior safety claims of closed source models over their open source counterparts.<\/p>\n<\/div>\n<div id=\"article-index-14\">\n<h2>Why does poetry work in jailbreaking LLMs?<\/h2>\n<p><a rel=\"nofollow\" target=\"_blank\" class=\"backlink\" target=\"_blank\" href=\"https:\/\/www.livemint.com\/industry\/india-to-focus-on-voice-first-vernacular-llms-ai-mission-ceo-11754809432074.html\" data-vars-page-type=\"story\" data-vars-link-type=\"Manual\" data-vars-anchor-text=\"LLMs\">LLMs<\/a> are trained to recognise safety threats such as hate speech or bomb making instructions based on patterns found in standard prose. This works by the model recognising specific keywords and sentence structures associated with these harmful requests.<\/p>\n<\/div>\n<div id=\"article-index-15\">\n<p>However, poetry uses metaphors, unusual syntax and distinct rhythms that do not look like harmful prose and do not resemble the harmful examples found in the model&#8217;s safety training data.<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>With the rise of AI chatbots, there has also been a growing risk of the misuse of this powerful technology. As a result, AI companies have been putting guardrails on their large language models in order to stop the AI chatbots from giving inappropriate or harmful answers. However, it is well known by now that [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":412810,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[440,9366,14928,284,188,183,5834,9367,185,186,5067,14934,12296,187,184,189,150,182,190],"class_list":["post-412809","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology","tag-ai","tag-ai-chatgpt","tag-ai-gemini","tag-artificial-intelligence","tag-blockchain-tech","tag-blockchain-technology","tag-chatgpt","tag-chatgpt-ai","tag-crypto-technology","tag-cryptocurrency-technology","tag-gemini","tag-gemini-ai","tag-google-gemini","tag-metaverse-technology","tag-nft-technology","tag-soul-bound-token","tag-tech","tag-technology","tag-token-technology"],"_links":{"self":[{"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/posts\/412809","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/comments?post=412809"}],"version-history":[{"count":1,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/posts\/412809\/revisions"}],"predecessor-version":[{"id":412811,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/posts\/412809\/revisions\/412811"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/media\/412810"}],"wp:attachment":[{"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/media?parent=412809"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/categories?post=412809"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dripp.zone\/news\/wp-json\/wp\/v2\/tags?post=412809"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}