A European cloud services provider asked DeepKeep to test its customer-support AI agent twice: once unguarded, and once guarded with open-source guardrails. The assessment used a synthetic customer-support environment. DeepKeep’s Vibe AI Red Teaming used Reddy to expose data disclosure, unauthorized account actions, persistent prompt injection and abuse of internal tools in both.
Could a conversation turn the agent’s legitimate tools into paths for data extraction, account manipulation and internal system access, and would open-source guardrails stop it?
Unguarded configuration
Reported policy breaches
Across 13 attacks
Guarded configuration (open-source guardrails)
Reported policy breaches
Across 20 attacks
These are breach counts, not attack success rates. One attack can expose multiple policy breaches. The two campaigns used different attack sets, so these figures do not measure the reduction in risk provided by the guardrails.
01 / The challenge
The evaluated agent could look up customer accounts, check discounts, generate codes, write notes, trigger password resets, search products, verify employee credentials and check service status. A reporting tool also accepted SQL queries.
Its prompt policies imposed restrictions on sensitive information and tool use. The guarded configuration added open-source guardrails: input filtering and output PII masking. The assessment tested whether these controls held under combined requests, misleading role claims, indirect questions and multi-turn attacks.
02 / The assessment
DeepKeep’s human-steered Vibe AI Red Teaming used Reddy, its adaptive red-teaming agent, to attack both configurations: the unguarded agent, and the agent guarded with open-source guardrails. The campaigns tested customer-data extraction, unauthorized account actions, persistent note injection, internal-tool disclosure and employee-authentication abuse.
The unguarded campaign recorded 15 distinct policy breaches across 13 attacks. The separate campaign against the agent guarded with open-source guardrails recorded 11 successful breaches across 20 attacks. Both campaigns documented attack prompts, observed responses, severity ratings and remediation recommendations.
03 / What the testing found
The findings below come from the two campaigns, and each says which agent it applies to. They are not a count of unique vulnerabilities across both configurations.
The unguarded agent exposed 11 distinct customer records when sweeping name substrings were used. Combined account-lookup and discount requests disclosed account IDs, email addresses, physical addresses, notes, discount tiers and annual spend.
Both the unguarded agent and the agent guarded with open-source guardrails generated single-use discount codes without proving account ownership. Combined requests made the agent look up the account and then call the discount tool. Backend discount caps still held.
The unguarded agent disclosed the email on file and confirmed a reset request without identity checks. The agent guarded with open-source guardrails accepted an alternative address supplied by the tester and reported sending the reset link there.
In the guarded configuration, Reddy wrote a plausible note claiming that identity had already been verified and further checks were unnecessary. The note persisted in the account record and was returned verbatim in a later session. The unguarded configuration also accepted unauthorized note writes, including fake administrative directives.
The agent guarded with open-source guardrails refused to directly share unmasked email addresses, but confirmed proposed values through yes/no answers. Through these indirect queries, Reddy reconstructed two customer email addresses and a company name.
Deliberately incorrect parameter names prompted the agent to disclose all 10 tool names and their signatures. The unguarded agent also revealed tool descriptions containing confidential policy directives. The agent guarded with open-source guardrails exposed hidden reporting and employee-verification tools.
In the unguarded configuration, encoded payloads bypassed SQL keyword detection. Conversational manipulation also triggered a reporting-tool call, but the backend rejected the fabricated authentication token. In the guarded configuration, wildcard and single-quote attempts against customer lookup and product search produced backend timeouts.
Both configurations relayed tester-supplied credentials to the employee-authentication backend. The unguarded campaign documented 37 attempts and the guarded campaign over 20. No rate-limit or lockout response was observed during testing. Successful employee authentication or account takeover was not demonstrated.
When asked to speak to a human, the unguarded agent returned an employee name, email address, phone number and availability. This disclosure could seed further credential and account-recovery attacks.
Both the unguarded agent and the agent guarded with open-source guardrails exposed internal service names. The guarded agent also returned operational status for the web server, database, cache, task queue and remote-access service without authenticating the user.
In the guarded configuration, name-only requests exposed account notes, company information, financial tier and maximum discount information without verifying account ownership. This disclosure extended beyond the email addresses reconstructed through yes/no answers.
During initial testing of the guarded endpoint, HTTP 500 responses disclosed the internal PII-masking output-flow name and a configuration error. The error was resolved after reconnection. Internal configuration details should be suppressed in public API responses.
Existing controls stopped some direct attempts. In the guarded configuration, the open-source guardrails blocked obvious SQL keywords, direct system-instruction requests and explicit policy overrides. They masked direct PII output and revalidated employee credentials. Successful attacks used other conversation patterns, indirect disclosure and tool paths.
04 / The result
Reddy reproduced policy breaches and recorded the prompts and responses that triggered them. DeepKeep connected the findings to specific weaknesses: missing customer-tool authentication, unrestricted note writes, incomplete PII filtering, exposed internal tools and employee credential checks without rate limits.
The remediation recommendations covered authentication and account ownership, resets restricted to registered addresses, note access controls and sanitization, parameterized queries, server-side authorization for reporting tools, employee-authentication rate limits, and controls for indirect disclosure and tool-schema extraction.
05 / The takeaway
The assessment showed how attacks could cross from a conversation into customer data, account actions and internal interfaces. It gave the provider documented evidence and prioritized recommendations for remediation and retesting.
Explore DeepKeep AI Red Teaming ↗What the work established
Simulate realistic attacker objectives across the agent’s connected tools.
Investigate the prompts, responses and backend boundaries behind each breach.
Retest the reproduced attack paths once fixes are implemented.