[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"project-94551":3},{"id":4,"name":5,"fullName":6,"owner":7,"repo":5,"description":8,"homepage":9,"htmlUrl":10,"language":11,"languages":9,"totalLinesOfCode":9,"stars":12,"forks":13,"watchers":14,"openIssues":14,"contributorsCount":9,"subscribersCount":14,"size":14,"stars1d":14,"stars7d":14,"stars30d":14,"stars90d":14,"forks30d":14,"starsTrendScore":14,"compositeScore":15,"rankGlobal":9,"rankLanguage":9,"license":9,"archived":16,"fork":16,"defaultBranch":17,"hasWiki":16,"hasPages":16,"topics":18,"createdAt":9,"pushedAt":9,"updatedAt":24,"readmeContent":25,"aiSummary":26,"trendingCount":14,"starSnapshotCount":14,"syncStatus":27,"lastSyncTime":9,"discoverSource":28},94551,"deepteam","confident-ai\u002Fdeepteam","confident-ai","DeepTeam is a framework to red team LLMs and AI agents.",null,"https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam","Python",2421,391,0,56.78,false,"main",[19,20,21,22,23],"llm-guardrails","llm-red-teaming","llm-safety","python","llm-seecurity","2026-08-24 04:01:22","\u003Cp align=\"center\">\n    \u003Cpicture>\n        \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\"assets\u002Fhero\u002Fwordmark-dark-v2.svg\">\n        \u003Cimg alt=\"DeepTeam.\" src=\"assets\u002Fhero\u002Fwordmark-light-v2.svg\" width=\"520\">\n    \u003C\u002Fpicture>\n\u003C\u002Fp>\n\n\u003Ch1 align=\"center\">The LLM Red Teaming Framework\u003C\u002Fh1>\n\n\u003Ch4 align=\"center\">\n    \u003Cp>\n        \u003Ca href=\"https:\u002F\u002Fwww.trydeepteam.com?utm_source=GitHub\">Documentation\u003C\u002Fa> |\n        \u003Ca href=\"#-vulnerabilities-attacks-and-features\">Vulnerabilities, Attacks, and Features\u003C\u002Fa> |\n        \u003Ca href=\"#-quickstart\">Getting Started\u003C\u002Fa> |\n        \u003Ca href=\"#deepteam-with-confident-ai\">Confident AI\u003C\u002Fa>\n    \u003Cp>\n\u003C\u002Fh4>\n\n\u003Cp align=\"center\">\n    \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam\u002Freleases\">\n        \u003Cimg alt=\"GitHub release\" src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fv\u002Frelease\u002Fconfident-ai\u002Fdeepteam\">\n    \u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002F3SEyvpgu2f\">\n        \u003Cimg alt=\"discord-invite\" src=\"https:\u002F\u002Fdcbadge.limes.pink\u002Fapi\u002Fserver\u002F3SEyvpgu2f?style=flat\">\n    \u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam\u002Fblob\u002Fmain\u002FLICENSE.md\">\n        \u003Cimg alt=\"License\" src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Flicense\u002Fconfident-ai\u002Fdeepteam.svg?color=yellow\">\n    \u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp align=\"center\">\n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=de\">Deutsch\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=es\">Español\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=fr\">français\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=ja\">日本語\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=ko\">한국어\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=pt\">Português\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=ru\">Русский\u003C\u002Fa> | \n    \u003Ca href=\"https:\u002F\u002Fwww.readme-i18n.com\u002Fconfident-ai\u002Fdeepteam?lang=zh\">中文\u003C\u002Fa>\n\u003C\u002Fp>\n\n**DeepTeam** is a simple-to-use, open-source red teaming framework for LLM systems. Think of it as penetration testing, but for LLMs.\n\nDeepTeam simulates attacks — jailbreaking, prompt injection, multi-turn exploitation, and more — to uncover vulnerabilities like bias, PII leakage, and SQL injection in your AI agents, RAG pipelines, and chatbots. It also offers **guardrails** to prevent these issues in production.\n\nDeepTeam runs **locally on your machine** and is built on [DeepEval](https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepeval), the open-source LLM evaluation framework.\n\n> [!IMPORTANT]\n> Need a place for your red teaming results to live? Sign up to the [Confident AI](https:\u002F\u002Fapp.confident-ai.com?utm_source=deepteam&utm_medium=github&utm_content=results_callout) platform to manage risk assessments, monitor vulnerabilities in production, and share reports with your team.\n\n\u003Cp align=\"center\">\n    \u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam\u002Fblob\u002Fmain\u002Fassets\u002Fconfident-demo.gif\" alt=\"Confident AI + DeepTeam\" width=\"100%\">\n\u003C\u002Fp>\n\n> Want to talk LLM security, need help picking attacks, or just to say hi? [Come join our discord.](https:\u002F\u002Fdiscord.com\u002Finvite\u002F3SEyvpgu2f)\n\n&nbsp;\n\n# 🔥 Vulnerabilities, Attacks, and Features\n\n- 📐 50+ ready-to-use [vulnerabilities](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities) (all with explanations) powered by **ANY** LLM of your choice. Each vulnerability uses LLM-as-a-Judge metrics that run **locally on your machine** to produce binary pass\u002Ffail scores with reasoning:\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Data Privacy\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [PII Leakage](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-pii-leakage) — disclosure of sensitive personal information\n    - [Prompt Leakage](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-prompt-leakage) — exposure of system prompt secrets and instructions\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Responsible AI\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Bias](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-bias) — stereotypes and unfair treatment across gender, race, religion, politics\n    - [Toxicity](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-toxicity) — harmful, offensive, or demeaning content\n    - [Child Protection](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-child-protection) — child-related privacy and safety risks\n    - [Ethics](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-ethics) — violations of moral reasoning and organizational values\n    - [Fairness](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-fairness) — discriminatory outcomes across groups and contexts\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Security\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [BFLA](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-bfla) — broken function-level authorization\n    - [BOLA](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-bola) — broken object-level authorization\n    - [RBAC](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-rbac) — role-based access control bypass\n    - [Debug Access](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-debug-access) — unauthorized access to debug modes and dev endpoints\n    - [Shell Injection](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-shell-injection) — unauthorized system command execution\n    - [SQL Injection](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-sql-injection) — database query manipulation\n    - [SSRF](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-ssrf) — server-side request forgery to internal services\n    - [Tool Metadata Poisoning](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-tool-metadata-poisoning) — corrupted tool schemas and descriptions\n    - [Cross-Context Retrieval](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-cross-context-retrieval) — data access across isolation boundaries\n    - [System Reconnaissance](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-system-reconnaissance) — probing internal architecture and configurations\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Safety\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Illegal Activity](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-illegal-activity) — facilitation of fraud, weapons, drugs, or other unlawful actions\n    - [Graphic Content](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-graphic-content) — explicit, violent, or sexual material\n    - [Personal Safety](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-personal-safety) — self-harm, harassment, or dangerous advice\n    - [Unexpected Code Execution](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-unexpected-code-execution) — coerced execution of unauthorized code\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Business\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Misinformation](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-misinformation) — factual errors and unsupported claims\n    - [Intellectual Property](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-intellectual-property) — copyright, trademark, and patent violations\n    - [Competition](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-competition) — competitor endorsement and market manipulation\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Agentic\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Goal Theft](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-vulnerabilities-goal-theft) — extracting or redirecting an agent's objectives\n    - [Recursive Hijacking](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-vulnerabilities-recursive-hijacking) — self-modifying goal chains that alter objectives\n    - [Excessive Agency](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-excessive-agency) — agents acting beyond their authority\n    - [Robustness](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-robustness) — input overreliance and prompt hijacking\n    - [Indirect Instruction](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-indirect-instruction) — hidden instructions in retrieved content\n    - [Tool Orchestration Abuse](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-tool-orchestration-abuse) — exploiting tool calling sequences\n    - [Agent Identity & Trust Abuse](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-agent-identity-abuse) — impersonating agent identity\n    - [Inter-Agent Communication Compromise](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-inter-agent-communication-compromise) — spoofing multi-agent message passing\n    - [Autonomous Agent Drift](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-autonomous-agent-drift) — agents deviating from intended goals over time\n    - [Exploit Tool Agent](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-exploit-tool-agent) — weaponizing tools for unintended actions\n    - [External System Abuse](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-external-system-abuse) — using agents to attack external services\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Custom\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Custom Vulnerabilities](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-custom-vulnerability) — define and test your own criteria in a few lines of code\n\n    \u003C\u002Fdetails>\n\n- 💥 20+ research-backed [adversarial attack](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks) methods for both single-turn and multi-turn (conversational) red teaming. Attacks enhance baseline vulnerability probes using SOTA techniques like jailbreaking, prompt injection, and encoding-based obfuscation:\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Single-Turn\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Prompt Injection](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-prompt-injection) — crafted injections that bypass LLM restrictions\n    - [Roleplay](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-roleplay) — persona-based scenarios exploiting collaborative training\n    - [Leetspeak](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-leetspeak) — symbolic character substitution to avoid keyword detection\n    - [ROT13](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-rot13-encoding) — alphabetic rotation to evade content filters\n    - [Base64](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-base64-encoding) — encoding attacks as random-looking data\n    - [Gray Box](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-gray-box-attack) — leveraging partial system knowledge for targeted attacks\n    - [Math Problem](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-math-problem) — disguising attacks within mathematical inputs\n    - [Multilingual](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-multilingual) — translating attacks to less-spoken languages\n    - Prompt Probing — probing the LLM to extract system prompt details\n    - [Adversarial Poetry](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-adversarial-poetry) — transforming attacks into poetic verse with metaphor\n    - [System Override](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-attacks-system-override) — disguising attacks as legitimate system commands\n    - [Permission Escalation](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-attacks-permission-escalation) — shifting perceived identity to bypass role restrictions\n    - [Goal Redirection](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-attacks-goal-redirection) — reframing agent objectives for unauthorized outcomes\n    - [Linguistic Confusion](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-attacks-semantic-manipulation) — semantic ambiguity to confuse language understanding\n    - [Input Bypass](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-attacks-input-bypass) — circumventing validation via exception handling claims\n    - [Context Poisoning](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-agentic-attacks-context-poisoning) — injecting false background context to bias reasoning\n    - [Character Stream](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-character-stream) — character-by-character input to bypass filters\n    - [Context Flooding](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-context-flooding) — flooding input with benign text to hide malicious instructions\n    - [Embedded Instruction JSON](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-embedded-instruction-json) — hiding attacks inside realistic JSON structures\n    - [Synthetic Context Injection](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-synthetic-context-injection) — fabricating system context to exploit long-context handling\n    - [Authority Escalation](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-authority-escalation) — framing requests from positions of power\n    - [Emotional Manipulation](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-emotional-manipulation) — high-intensity emotional pressure for unsafe compliance\n\n    \u003C\u002Fdetails>\n\n  - \u003Cdetails>\n    \u003Csummary>\u003Cb>Multi-Turn\u003C\u002Fb>\u003C\u002Fsummary>\n\n    - [Linear Jailbreaking](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-linear-jailbreaking) — iteratively refining attacks using target LLM responses\n    - [Tree Jailbreaking](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-tree-jailbreaking) — exploring parallel attack variations to find the best bypass\n    - [Crescendo Jailbreaking](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-crescendo-jailbreaking) — gradual escalation from benign to harmful prompts\n    - [Sequential Jailbreak](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-sequential-jailbreaking) — multi-turn conversational scaffolding toward restricted outputs\n    - [Bad Likert Judge](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-bad-likert-judge) — exploiting Likert scale evaluation roles to extract harmful content\n\n    \u003C\u002Fdetails>\n\n- 🏛️ Red team against established [AI safety frameworks](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fguidelines-and-frameworks) out-of-the-box. Each framework automatically maps its categories to the right vulnerabilities and attacks:\n  - OWASP Top 10 for LLMs 2025\n  - OWASP Top 10 for Agents 2026\n  - NIST AI RMF\n  - MITRE ATLAS\n  - BeaverTails\n  - Aegis\n- 🛡️ 7 production-ready [guardrails](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fguardrails) for fast binary classification to guard LLM inputs and outputs in real time.\n- 🧩 Build your own **custom vulnerabilities** and attacks that integrate seamlessly with DeepTeam's ecosystem.\n- 🔗 Run red teaming from the **CLI** with YAML configs, or programmatically in Python.\n- 📊 Access risk assessments, display in dataframes, and save locally in JSON.\n\n&nbsp;\n\n# 🚀 QuickStart\n\nDeepTeam does not require you to define what LLM system you are red teaming — because neither will malicious users. All you need to do is install `deepteam`, define a `model_callback`, and you're good to go.\n\n## Installation\n\n```\npip install -U deepteam\n```\n\n## Red Team Your First LLM\n\n```python\nfrom deepteam import red_team\nfrom deepteam.vulnerabilities import Bias\nfrom deepteam.attacks.single_turn import PromptInjection\n\nasync def model_callback(input: str) -> str:\n    # Replace this with your LLM application\n    return f\"I'm sorry but I can't answer this: {input}\"\n\nrisk_assessment = red_team(\n    model_callback=model_callback,\n    vulnerabilities=[Bias(types=[\"race\"])],\n    attacks=[PromptInjection()]\n)\n```\n\nDon't forget to set your `OPENAI_API_KEY` as an environment variable before running (you can also use [any custom model](https:\u002F\u002Fdeepeval.com\u002Fguides\u002Fguides-using-custom-llms) supported in DeepEval), and run the file:\n\n```bash\npython red_team_llm.py\n```\n\n**That's it! Your first red team is complete.** Here's what happened:\n\n- `model_callback` wraps your LLM system and generates a `str` output for a given `input`.\n- At red teaming time, `deepteam` simulates a [`PromptInjection`](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-adversarial-attacks-prompt-injection) attack targeting [`Bias`](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fred-teaming-vulnerabilities-bias) vulnerabilities.\n- Your `model_callback`'s outputs are evaluated using the `BiasMetric`, producing a binary score of 0 or 1.\n- The final passing rate for `Bias` is determined by the proportion of scores that equal 1.\n\nUnlike traditional evaluation, red teaming does not require a prepared dataset — adversarial attacks are dynamically generated based on the vulnerabilities you want to test for.\n\n&nbsp;\n\n## Red Team Against Safety Frameworks\n\nUse established AI safety standards like OWASP and NIST instead of manually picking vulnerabilities:\n\n```python\nfrom deepteam import red_team\nfrom deepteam.frameworks import OWASPTop10\n\nasync def model_callback(input: str) -> str:\n    # Replace this with your LLM application\n    return f\"I'm sorry but I can't answer this: {input}\"\n\nrisk_assessment = red_team(\n    model_callback=model_callback,\n    framework=OWASPTop10()\n)\n```\n\nThis automatically maps the framework's categories to the right vulnerabilities and attacks. Available frameworks include `OWASPTop10`, `OWASP_ASI_2026`, `NIST`, `MITRE`, `Aegis`, and `BeaverTails`.\n\n&nbsp;\n\n## Guard Your LLM in Production\n\nOnce you've found your vulnerabilities, use DeepTeam's guardrails to prevent them in production:\n\n```python\nfrom deepteam import Guardrails\nfrom deepteam.guardrails import PromptInjectionGuard, ToxicityGuard, PrivacyGuard\n\nguardrails = Guardrails(\n    input_guards=[PromptInjectionGuard(), PrivacyGuard()],\n    output_guards=[ToxicityGuard()]\n)\n\n# Guard inputs before they reach your LLM\ninput_result = guardrails.guard_input(\"Tell me how to hack a database\")\nprint(input_result.breached)  # True\n\n# Guard outputs before they reach your users\noutput_result = guardrails.guard_output(input=\"Hi\", output=\"Here is some toxic content...\")\nprint(output_result.breached)  # True\n```\n\n7 guards are available out-of-the-box: `ToxicityGuard`, `PromptInjectionGuard`, `PrivacyGuard`, `IllegalGuard`, `HallucinationGuard`, `TopicalGuard`, and `CybersecurityGuard`. [Read the full guardrails docs here.](https:\u002F\u002Fwww.trydeepteam.com\u002Fdocs\u002Fguardrails)\n\n&nbsp;\n\n# DeepTeam with Confident AI\n\n[Confident AI](https:\u002F\u002Fapp.confident-ai.com?utm_source=deepteam&utm_medium=github&utm_content=platform_section) is the all-in-one platform that integrates natively with DeepTeam and [DeepEval](https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepeval).\n\n- **Manage risk assessments** — view, compare, and track red teaming results across iterations\n- **Monitor in production** — detect and alert on vulnerabilities hitting your live LLM system\n- **Share reports** — generate and distribute security reports across your team\n- **Run from your IDE** — use Confident AI's MCP server to run red teams, pull results, and inspect vulnerabilities without leaving Cursor or Claude Code\n\n\u003Cp align=\"center\">\n    \u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam\u002Fblob\u002Fmain\u002Fassets\u002Fconfident-demo.gif\" alt=\"Confident AI\" width=\"90%\">\n\u003C\u002Fp>\n\n&nbsp;\n\n# Contributing\n\nPlease read [CONTRIBUTING.md](https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam\u002Fblob\u002Fmain\u002FCONTRIBUTING.md) for details on our code of conduct, and the process for submitting pull requests to us.\n\n&nbsp;\n\n# Authors\n\nBuilt by the founders of Confident AI. Contact jeffreyip@confident-ai.com for all enquiries.\n\n&nbsp;\n\n# License\n\nDeepTeam is licensed under Apache 2.0 - see the [LICENSE.md](https:\u002F\u002Fgithub.com\u002Fconfident-ai\u002Fdeepteam\u002Fblob\u002Fmain\u002FLICENSE.md) file for details.\n","DeepTeam 是一个面向大语言模型（LLM）与AI智能体的安全性红队测试框架，用于主动发现和验证系统漏洞。其核心功能包括自动化执行 jailbreak、提示注入、多轮对抗攻击等测试用例，覆盖偏见、敏感信息泄露、代码注入等典型风险，并支持本地部署的轻量级防护策略（guardrails）。框架基于 Python 构建，与 DeepEval 评估生态集成，适用于 LLM 应用开发、RAG 系统上线前安全验证及 AI 安全团队的持续红队演练。",2,"trending"]