Hands-on Practice
student workspace
0 XP
PS
AI Foundation Program/AI Security and Ethics
intermediate7 min read

Cybersecurity and AI Threats

AI is a security double-edged sword: it strengthens defenses, but it also introduces entirely new categories of attack. Here's what to watch for.

Prompt injection

Malicious instructions hidden in content an AI reads, tricking it into ignoring its original task.

Data poisoning

Attackers corrupt training data so a model learns harmful or biased behavior.

Model theft

Competitors or attackers extract a model's behavior or weights through repeated querying.

AI-powered phishing

Generative AI writes highly convincing, personalized phishing emails at scale.

Adversarial examples

Inputs subtly altered to fool a model — like a stop sign with stickers that a self-driving car misreads.

Practical defensive practices

  1. 1

    Treat AI outputs as untrusted input

    Never let an AI system take irreversible action without a human or rule-based check.

  2. 2

    Sanitize content fed to AI agents

    Strip or flag suspicious instructions in documents, emails, or web pages before an AI reads them.

  3. 3

    Limit blast radius

    Give AI agents the minimum permissions and tool access they actually need.

  4. 4

    Monitor and log

    Keep an audit trail of AI-driven actions so anomalies can be caught and traced.

Key takeaways

  • Prompt injection, data poisoning, model theft, AI-powered phishing, and adversarial examples are distinct new AI-era threats.
  • AI outputs and AI-read content should be treated as untrusted input, not automatically acted upon.
  • Limiting an AI agent's permissions and logging its actions reduces the impact of any single compromise.

Check your understanding

0/2 answered

1.What is 'prompt injection'?

2.It's good practice to let AI agents take irreversible actions automatically, without any human or rule-based check.

Lesson summary

AI introduces new attack surfaces — prompt injection, data poisoning, model theft, and more — that require treating AI inputs/outputs as untrusted and limiting agent permissions.

AI-generated notes