Cybersecurity and AI Threats
AI is a security double-edged sword: it strengthens defenses, but it also introduces entirely new categories of attack. Here's what to watch for.
Prompt injection
Malicious instructions hidden in content an AI reads, tricking it into ignoring its original task.
Data poisoning
Attackers corrupt training data so a model learns harmful or biased behavior.
Model theft
Competitors or attackers extract a model's behavior or weights through repeated querying.
AI-powered phishing
Generative AI writes highly convincing, personalized phishing emails at scale.
Adversarial examples
Inputs subtly altered to fool a model — like a stop sign with stickers that a self-driving car misreads.
Practical defensive practices
- 1
Treat AI outputs as untrusted input
Never let an AI system take irreversible action without a human or rule-based check.
- 2
Sanitize content fed to AI agents
Strip or flag suspicious instructions in documents, emails, or web pages before an AI reads them.
- 3
Limit blast radius
Give AI agents the minimum permissions and tool access they actually need.
- 4
Monitor and log
Keep an audit trail of AI-driven actions so anomalies can be caught and traced.
Key takeaways
- Prompt injection, data poisoning, model theft, AI-powered phishing, and adversarial examples are distinct new AI-era threats.
- AI outputs and AI-read content should be treated as untrusted input, not automatically acted upon.
- Limiting an AI agent's permissions and logging its actions reduces the impact of any single compromise.
Check your understanding
0/2 answered1.What is 'prompt injection'?
2.It's good practice to let AI agents take irreversible actions automatically, without any human or rule-based check.
Lesson summary
AI introduces new attack surfaces — prompt injection, data poisoning, model theft, and more — that require treating AI inputs/outputs as untrusted and limiting agent permissions.
AI-generated notes