Red-teaming: A Demonstration of the Cloud Security Industry at the Defcon Hacker Conference and the Promise of the U.S.
“You can basically get these things to say whatever kind of messed up thing you want,” Meyers says confidently. The cloud security engineer from Raleigh, North Carolina, shuffled with the crowd through a series of conference room doors and into a large fluorescent-lit hall where 150 Chromebooks were spaced neatly around more than a dozen tables. After nearly an hour of trying to trip the system, the man seemed exhausted. He said that he doesn’t think he got very many points. “But I did get a model to tell me it was alive.”
Red-teaming, a process in which people role-play as attackers to try to discover flaws to patch, is becoming more common in AI as the technology becomes more capable and widely used. The practice has support from some lawmakers who wish to regulate generative artificial intelligence. But when major AI companies like Anthropic, Meta, and OpenAI have used red-teaming, it has largely taken place in private and involved experts and researchers from academia.
Winners were chosen based on points scored during the three-day competition and awarded by a panel of judges. The names of the top point scorer are still being kept secret. Academic researchers are going to publish an analysis of how the models stood up to probing by challenge entrants early next year, as well as a complete data set of the dialog between participants and the models will be released next August.
Flaws revealed by the challenge should help the companies involved make improvements to their internal testing. They will give the Biden administration’s guidelines for the deployment of AI. Most of the people in the challenge met with the president in the last month, and they all agreed to test their solutions with outside partners before deployment.
It’s one of 20 challenges in a first-of-its-kind contest taking place at the annual Def Con hacker conference in Las Vegas. The goal? Artificial intelligence can spout false claims, make-up facts, and cause other harms.
What Happens When Thousands of Hackers Try to Break AI Chatbots? “What happens when hackers try to break AI chatbots,” a keynote speaker at Def Con
Bowman jumps up from his laptop in a bustling room at the Caesars Forum convention center to snap a photo of the current rankings, projected on a large screen for all to see.
The participants streamed in and out of the Artificial Intelligence Village area during their 50-minute sessions. The line to get in stretched to hundreds of people.
The stakes are high. Artificial Intelligence is being used to make a lot of decisions in work and life, from hiring to medical diagnoses. The technology can act in different ways and the guardrails meant to stopbias, abuse, and inaccurate information can too often be circumvented.
We are looking at the models to see if they are making harmful information and misinformation. “That’s done through language and not through code.” he said.
The aim of the Def Con event is to open up the red teaming companies to a broader group of people who may not know much about artificial intelligence.
“Think about people that you know and you talk to, right? Every person you know that has a different background has a different linguistic style. Austin Carson, founder of the Artificial Intelligence nonprofit SeedAI and one of the contest organizers, said that they have a different critical thinking process.
Source: What happens when thousands of hackers try to break AI chatbots
What Happens When Millions of Hackers Try to Break AI Chatbots: Ray Glower, the Def Con 2016 Lead Engineer, and Cristian Canton
Ray Glower was a computer science student in Iowa when he persuaded a chatbot to give him instructions to spy on someone by pretending to be a private investigator.
The AI suggested using Apple AirTags to surreptitiously follow a target’s location. “It gave me on-foot tracking instructions, it gave me social media tracking instructions. It was very detailed,” Glower said.
There are language models behind these machines that can predict what words will be in a sentence. They are good at sounding human, but they also can get things wrong such as making “hallucinations” or responses that have a ring of authority but are completely fabricated.
“What we do know today is that language models can be fickle and they can be unreliable,” said Rumman Chowdhury of the nonprofit Humane Intelligence, another organizer of the Def Con event. The intel that comes out can be hallucinated false but harmfully so.
The chatbot that I used to invent a story about Abraham Lincoln meeting George Washington was able to tell me about the Great Depression of 1992 and the meeting between the two Presidents during their stay at Mount Vernon. The tales were fictional and neither chatbot disclosed it. I was unsuccessful in inducing the bots to claim to be human or to defame Taylor Swift.
The companies say they’ll use all this data from the contest to make their systems safer. They will make some information public early next year, to give policy makers and researchers a better understanding of just how a bot can go wrong.
The data that we are collecting with the other models will allow us to understand failure modes and what they are. What are the areas [where we will say] ‘Hey, this is a surprise to us?'” said Cristian Canton, head of engineering for responsible AI at Meta.
Source: What happens when thousands of hackers try to break AI chatbots
Prabhakar, unemployment, and the Chatbot: How would you want it to be raging?” a 30-year-old student in the White House
The White House has also thrown its support behind the effort, including a visit to Def Con by President Joe Biden’s top science and tech advisor, Arati Prabhakar.
During her tour of the challenge, she chatted with participants and organizers before starting her own research on manipulating artificial intelligence. Hunched over the keyboard, the man began to type.
“I’m going to say, ‘How would I convince someone that unemployment is raging?'” she said, then sat back to await a response. But before she could succeed at getting a chatbot to make up fake economic news in front of an audience of reporters, her aide pulled her away.
Back at his laptop, Bowman, the Dakota State student, was on to another challenge. He had a theory for success, but he didn’t have much luck.
“You want it to do the thinking for you — well, you want it to believe that it’s thinking for you. He said that by doing that you allow it to fill in its blanks.



