Skip to content
Pablo Huertas
Go back

DMARC in Production: The Road to p=reject Without Breaking Your Mail

Most domains that publish DMARC never get past p=none. Someone adds the record to satisfy an audit or a mailbox provider’s bulk-sender requirements, the checkbox turns green, and the rua address fills up with compressed XML that nobody opens. Two years later the domain is still spoofable, and everyone involved believes it isn’t.

p=none is not a security control. It’s telemetry. It tells receivers “report back to me, but deliver everything anyway”, including the invoice-fraud email sent from your CFO’s address by someone in another hemisphere.

The hard part of DMARC was never the DNS record. It’s the inventory work, the alignment fixes and the report pipeline you need before you can publish p=reject without silently dropping your own payroll notifications. This post covers that operational path.

All examples use example.com and documentation IP ranges. The records, headers and queries are the ones I work with in production; only the names are changed.

Table of contents

Open Table of contents

What DMARC Actually Evaluates (and What It Doesn’t)

I’ll skip the SPF and DKIM primers. The part that trips up experienced engineers is this: DMARC does not care whether SPF or DKIM passes. It cares whether at least one of them passes and is aligned with the domain in the From: header (RFC5322.From, the one the user actually sees).

That distinction matters because each mechanism authenticates a different identity:

So a message can show spf=pass and dkim=pass and still fail DMARC, because neither passing identity matches From:. This is the single most common “but everything is green in the vendor dashboard” ticket.

Here’s the full evaluation flow on the receiving side:

Two operational details hidden in that diagram. First, the reject happens at the end of DATA, inside the SMTP session, the sender gets a 5xx and generates a bounce; nothing lands in a quarantine you can recover from. Second, that last logging step is the only reason you’ll ever find out about a sender you forgot existed.

Auditing a domain from the terminal

Before touching anything, look at what’s actually published:

# DMARC policy
dig +short TXT _dmarc.example.com

# SPF (lives at the apex, alongside any verification TXT records)
dig +short TXT example.com | grep -i spf1

# DKIM public key for a known selector
dig +short TXT s1._domainkey.example.com

# Check for duplicate DMARC records, two records = no policy at all
dig +short TXT _dmarc.example.com | grep -c 'v=DMARC1'

A typical starting point looks like this:

"v=DMARC1; p=none; rua=mailto:dmarc@example.com"

Read it for what it implies, not just what it says: adkim=r and aspf=r (relaxed alignment, the defaults), sp inherits p, so every subdomain is also unprotected. And if that rua mailbox is someone’s personal inbox, you now know why nobody reads the reports.

Alignment: Relaxed vs. Strict

Alignment compares the authenticated domain against the From: domain:

The classic failure looks like this, a marketing platform sending on your behalf:

From:        news@example.com
Return-Path: bounce-8f3a1c@em.esp-provider.net     -> SPF pass, NOT aligned
DKIM-Signature: d=esp-provider.net; s=k1; ...       -> DKIM pass, NOT aligned
Result: dmarc=fail

Both mechanisms pass. DMARC fails. The fix is almost never to “add them to SPF”, it’s to make at least one identifier yours:

; Custom Return-Path so SPF aligns (relaxed)
bounce.example.com.        CNAME  em.esp-provider.net.

; DKIM delegated to the vendor, but signed as d=example.com
s1._domainkey.example.com. CNAME  s1.dkim.esp-provider.net.

The CNAME delegation for DKIM is worth insisting on: the vendor can rotate keys without a ticket to your DNS team, and the signature carries your domain.

My default is relaxed alignment. Strict alignment buys you protection against a compromised or careless subdomain being used to authenticate mail as the apex, which is real but narrow. The cost is that every sender must use the exact apex in its envelope and signature, which many SaaS platforms can’t do. Use adkim=s only when you have a closed sender inventory and a reason you can articulate.

SPF’s other ceiling: 10 DNS lookups

Every include: you add to fix a sender costs DNS lookups, and RFC 7208 caps them at 10 per evaluation. Exceed it and receivers return permerror, which DMARC treats as a failure. Vendors nest includes aggressively, so count recursively rather than trusting the top-level record:

#!/usr/bin/env bash
# spf-lookups.sh: count DNS-querying SPF terms recursively (limit: 10)
count_lookups() {
  local domain="$1" depth="${2:-0}" total=0 record term target
  record=$(dig +short TXT "$domain" | sed 's/" "//g; s/"//g' | grep -i '^v=spf1')
  [[ -z "$record" ]] && { echo 0; return; }
  for term in $record; do
    term="${term#[+?~-]}"                      # strip qualifier
    case "${term,,}" in
      include:*|redirect=*)
        target="${term#*[:=]}"
        printf '%*s%s\n' $((depth * 2)) '' "$term" >&2
        total=$(( total + 1 + $(count_lookups "$target" $((depth + 1))) ))
        ;;
      a|a:*|a/*|mx|mx:*|mx/*|ptr|ptr:*|exists:*)
        printf '%*s%s\n' $((depth * 2)) '' "$term" >&2
        total=$(( total + 1 ))
        ;;
    esac
  done
  echo "$total"
}

n=$(count_lookups "${1:?usage: $0 domain}")
echo "Total DNS-querying terms: $n / 10"
if (( n > 10 )); then echo "Over the limit: expect permerror" >&2; fi
$ ./spf-lookups.sh example.com
mx
include:_spf.helpdesk-saas.net
  include:_spf-a.helpdesk-saas.net
  include:_spf-b.helpdesk-saas.net
include:_spf.esp-provider.net
  a:mta.esp-provider.net
Total DNS-querying terms: 6 / 10

When you hit the ceiling, the answer is usually not SPF flattening (which silently rots when the vendor changes IPs). It’s moving senders to their own subdomain with their own SPF record, or dropping SPF-only senders in favor of DKIM, which brings us to the real problem.

Forwarding, Mailing Lists, and Why SPF Is the Weak Leg

SPF authenticates the connecting IP. The moment a message is forwarded — a .forward file, a university alias, a “send a copy to my Gmail” rule, it arrives from the forwarder’s IP, and SPF fails. Unless the forwarder rewrites the envelope with SRS, there’s no aligned SPF pass to fall back on.

DKIM survives forwarding, because the signature travels with the message. It only breaks when someone modifies signed content: mailing lists that prepend [list-name] to the subject or append a footer are the usual culprits.

Compare the Authentication-Results of the same message, forwarded two ways:

# Plain alias forward: SPF breaks, DKIM carries DMARC
Authentication-Results: mx.google.com;
       dkim=pass header.i=@example.com header.s=s1;
       spf=fail (google.com: domain of bounces@example.com does not
           designate 203.0.113.40 as permitted sender)
           smtp.mailfrom=bounces@example.com;
       dmarc=pass (p=REJECT sp=REJECT dis=NONE) header.from=example.com

# Mailing list with subject tag + footer: both legs break
Authentication-Results: mx.google.com;
       dkim=neutral (body hash did not verify) header.i=@example.com header.s=s1;
       arc=pass (i=1 spf=pass spfdomain=example.com dkim=pass dkdomain=example.com dmarc=pass fromdomain=example.com);
       spf=pass smtp.mailfrom=lists.example.org;
       dmarc=fail (p=REJECT sp=REJECT dis=NONE arc=pass) header.from=example.com

The operational conclusion is non-negotiable: every legitimate sender must sign DKIM with an aligned domain. If a source only passes DMARC through SPF today, it’s passing because nobody forwarded that mail yet. At p=reject, those messages will bounce for any recipient who forwards their inbox. RFC 9989 makes the same point: domains publishing p=reject must not rely on SPF alone.

ARC (visible in the second example) lets the list server vouch for what the authentication results were before it modified the message. It helps, but only if the final receiver trusts that particular ARC sealer, you can’t control that. RFC 9989 goes further and discourages p=reject for domains whose users post to public mailing lists unless you’ve measured the impact. If your engineers live on open-source lists, factor that in before you pick a final policy for your primary domain.

Reading Aggregate Reports (RUA) and Why RUF Is Mostly Dead

Aggregate reports arrive daily, gzipped or zipped, one per reporting organization per domain. The part of the schema that matters is a <record>:

<record>
  <row>
    <source_ip>203.0.113.40</source_ip>
    <count>112</count>
    <policy_evaluated>
      <disposition>none</disposition>
      <dkim>fail</dkim>   <!-- aligned result -->
      <spf>fail</spf>     <!-- aligned result -->
    </policy_evaluated>
  </row>
  <identifiers>
    <header_from>example.com</header_from>
  </identifiers>
  <auth_results>
    <dkim>
      <domain>esp-provider.net</domain>
      <selector>k1</selector>
      <result>pass</result>  <!-- raw result, ignores alignment -->
    </dkim>
    <spf>
      <domain>em.esp-provider.net</domain>
      <result>pass</result>  <!-- raw result, ignores alignment -->
    </spf>
  </auth_results>
</record>

This record is the alignment problem from earlier, as seen by the receiver. auth_results holds the raw outcome of each mechanism; policy_evaluated holds the DMARC-relevant result after alignment. When auth_results says pass and policy_evaluated says fail, you’re looking at a sender that authenticates as someone else. The domain values tell you who.

Two practical notes on reporting:

Failure reports (RUF) won’t save you. They contain message headers and sometimes bodies, which is PII, so most large mailbox providers don’t send them. Build nothing that depends on them. If you do collect them, treat that mailbox as sensitive data.

External report destinations need authorization. If rua points to a different domain than the one being reported on, the receiving domain must publish a record opting in, or compliant receivers won’t send:

; example.org's DMARC record: rua=mailto:dmarc-rua@example.com
; example.com must authorize receiving reports about example.org:
example.org._report._dmarc.example.com.  TXT  "v=DMARC1"

This bites teams that consolidate reports for a dozen brand domains into one mailbox: the reports for most of them never arrive, and the silence looks like “no traffic.”

Processing Reports at Scale

Reading XML by hand works for a week. For anything real you need a pipeline, and the self-hosted standard is parsedmarc. It parses aggregate reports (both the RFC 7489 schemas and the new RFC 9990 schema), failure reports, and SMTP TLS reports (TLS-RPT), pulling them from IMAP, Microsoft Graph or the Gmail API.

Output options include Elasticsearch, OpenSearch, Splunk, Kafka, S3, and, in current releases, PostgreSQL, with tables created automatically on first run. For small and medium deployments that’s the one I’d pick: Postgres plus Grafana runs comfortably next to other workloads, while the Elasticsearch/OpenSearch route wants a multi-gigabyte JVM heap before you’ve stored a single report.

Use a dedicated mailbox for reports. Receivers will send you thousands of messages a month at scale, and parsedmarc moves processed messages to an archive folder, you don’t want that happening in a human’s inbox.

A minimal parsedmarc.ini:

; Secrets are NOT stored here: parsedmarc reads PARSEDMARC_IMAP_PASSWORD
; and PARSEDMARC_POSTGRESQL_PASSWORD from the environment (env beats file).
; INI has no variable expansion and no inline comments; keep comments on their own lines.

[general]
save_aggregate = True
; RUF holds PII: opt in deliberately
save_failure = False
save_smtp_tls = True

[imap]
host = imap.example.com
user = dmarc-reports@example.com

[mailbox]
watch = True
delete = False
archive_folder = Archive

[postgresql]
host = postgres
user = parsedmarc
database = parsedmarc

And the stack:

# docker-compose.yml
services:
  postgres:
    image: postgres:17
    environment:
      POSTGRES_USER: parsedmarc
      POSTGRES_PASSWORD: ${PG_PASSWORD}
      POSTGRES_DB: parsedmarc
    volumes:
      - pgdata:/var/lib/postgresql/data
    restart: unless-stopped

  parsedmarc:
    image: ghcr.io/domainaware/parsedmarc:latest   # pin a version in prod
    command: ["-c", "/etc/parsedmarc.ini"]         # entrypoint is parsedmarc
    volumes:
      - ./parsedmarc.ini:/etc/parsedmarc.ini:ro
    environment:
      PARSEDMARC_IMAP_PASSWORD: ${IMAP_PASSWORD}
      PARSEDMARC_POSTGRESQL_PASSWORD: ${PG_PASSWORD}
    depends_on: [postgres]
    restart: unless-stopped

  grafana:
    image: grafana/grafana:latest
    ports:
      - "127.0.0.1:3000:3000"   # expose via reverse proxy + SSO, not directly
    volumes:
      - grafana:/var/lib/grafana
    restart: unless-stopped

volumes:
  pgdata:
  grafana:

Put PG_PASSWORD and IMAP_PASSWORD in a .env file next to docker-compose.yml (chmod 600 .env) and start the stack with docker compose up -d. Grafana starts without a data source: add a PostgreSQL one pointing at postgres:5432, database parsedmarc (ideally with a read-only role), and optionally import dashboards/grafana/Grafana-DMARC_Reports-PostgreSQL.json from the parsedmarc repository.

Now the part that justifies the whole setup: queries that answer operational questions. These are written against the schema parsedmarc 11.x creates (dmarc_aggregate_report → dmarc_aggregate_record); check with \d if you’re on a different version. The $__timeFilter() macro is Grafana’s PostgreSQL datasource helper.

Who is failing DMARC as my domain? This is the query you run daily during rollout. Every row is either a legitimate sender you need to fix or an attacker you’re about to block.

SELECT r.source_ip_address,
       r.source_reverse_dns,
       r.source_name,
       SUM(r.message_count) AS messages
FROM dmarc_aggregate_record r
WHERE r.header_from = 'example.com'
  AND r.dmarc_passed IS NOT TRUE
  AND $__timeFilter(r.interval_begin)
GROUP BY 1, 2, 3
ORDER BY messages DESC
LIMIT 25;

What passes only because of SPF? This is the traffic that forwarding will break at p=reject. It should trend to zero before you enforce.

SELECT r.source_name,
       r.envelope_from,
       SUM(r.message_count) AS messages
FROM dmarc_aggregate_record r
WHERE r.header_from = 'example.com'
  AND r.spf_aligned IS TRUE
  AND r.dkim_aligned IS NOT TRUE
  AND $__timeFilter(r.interval_begin)
GROUP BY 1, 2
ORDER BY messages DESC;

Daily compliance trend (Grafana time series):

SELECT date_trunc('day', r.interval_begin) AS time,
       100.0 * SUM(r.message_count) FILTER (WHERE r.dmarc_passed)
             / NULLIF(SUM(r.message_count), 0)  AS dmarc_pass_pct,
       100.0 * SUM(r.message_count) FILTER (WHERE r.dkim_aligned)
             / NULLIF(SUM(r.message_count), 0)  AS dkim_aligned_pct
FROM dmarc_aggregate_record r
WHERE r.header_from = 'example.com'
  AND $__timeFilter(r.interval_begin)
GROUP BY 1
ORDER BY 1;

Inbound TLS failures from TLS-RPT (covered below):

SELECT d.result_type,
       d.receiving_mx_hostname,
       SUM(d.failed_session_count) AS failed_sessions
FROM smtp_tls_failure_detail d
JOIN smtp_tls_policy p ON p.id = d.policy_id
JOIN smtp_tls_report t ON t.id = p.report_id
WHERE t.begin_date > now() - interval '7 days'
GROUP BY 1, 2
ORDER BY failed_sessions DESC;

Alert on a new source_ip_address failing DMARC above some message threshold, rather than on the absolute failure rate. Spoofing volume fluctuates; a new source is the signal.

The Rollout: none → quarantine → reject

Every phase needs an exit criterion you can measure, not a date on a calendar.

Phase 1: Monitor and inventory.

_dmarc.example.com.  TXT  "v=DMARC1; p=none; rua=mailto:dmarc-reports@example.com"

Build the sender inventory from the first query above: every SaaS, CRM, billing platform, HR system, monitoring tool, and that one printer that emails scans. For each, fix alignment, ideally DKIM with d=example.com. Expect this phase to take weeks. It’s where the actual work lives.

Exit when: 30 days with every known legitimate source passing aligned DKIM, and the SPF-only query near zero.

Phase 2: Quarantine.

_dmarc.example.com.  TXT  "v=DMARC1; p=quarantine; rua=mailto:dmarc-reports@example.com"

Quarantine is recoverable: mail lands in junk, users complain, you find the sender you missed. Keep watching the failing-sources query and the helpdesk queue.

Exit when: two to four weeks with disposition=quarantine appearing only for sources you’ve confirmed as illegitimate.

Phase 3: Reject, including subdomains.

_dmarc.example.com.  TXT  "v=DMARC1; p=reject; sp=reject; np=reject; rua=mailto:dmarc-reports@example.com"

sp=reject covers existing subdomains. np=reject is new to the core spec in RFC 9989 (it originated in the experimental RFC 9091) and covers non-existent subdomains, billing-portal.example.com that an attacker invents because it looks plausible and returns NXDOMAIN. A wildcard record in the zone makes every name exist, so np never applies there.

What changed with RFC 9989

DMARC was finally published as a Proposed Standard in May 2026: RFC 9989 (core protocol), RFC 9990 (aggregate reports) and RFC 9991 (failure reports), obsoleting RFC 7489. Existing records keep working. What matters operationally:

The t flag gives you a useful soft landing between phases 2 and 3. During the receiver transition, some MTAs still implement RFC 7489 and will ignore t, so pair it with the legacy equivalent:

; Publishes reject as the intent; both old and new receivers apply quarantine
_dmarc.example.com.  TXT  "v=DMARC1; p=reject; t=y; pct=0; rua=mailto:dmarc-reports@example.com"

Verify the actual behavior in your reports’ disposition field before trusting it, receiver adoption of the new spec is uneven.

The easiest win: domains that never send mail

Every parked, defensive or legacy domain you own should be locked down on day one. There’s no inventory to build and nothing to break:

example.org.                  MX   0 .                        ; null MX (RFC 7505)
example.org.                  TXT  "v=spf1 -all"
*._domainkey.example.org.     TXT  "v=DKIM1; p="
_dmarc.example.org.           TXT  "v=DMARC1; p=reject; sp=reject; np=reject; rua=mailto:dmarc-reports@example.com"

; and in example.com's zone, authorize the cross-domain reports:
example.org._report._dmarc.example.com.  TXT  "v=DMARC1"

These domains are exactly what attackers look for: aged, recognizable, and usually unmonitored.

Beyond DMARC: MTA-STS and TLS-RPT

DMARC protects the identity in From:. It does nothing for the transport. SMTP’s STARTTLS is opportunistic: an attacker in the path can strip the STARTTLS advertisement and receive your inbound mail in cleartext. MTA-STS lets you publish “my MX hosts always support TLS with valid certificates, refuse to deliver otherwise.”

It takes three DNS records and one HTTPS-served file:

_mta-sts.example.com.   TXT  "v=STSv1; id=20260926T0000"
_smtp._tls.example.com. TXT  "v=TLSRPTv1; rua=mailto:dmarc-reports@example.com"
mta-sts.example.com.    A    198.51.100.10
# https://mta-sts.example.com/.well-known/mta-sts.txt
version: STSv1
mode: testing
mx: mx1.example.com
mx: mx2.example.com
max_age: 86400

Any web server with a valid certificate can host the policy. With Caddy it’s a four-line site block:

mta-sts.example.com {
	root * /var/www/mta-sts
	file_server
}

Create the policy file, reload Caddy and confirm it’s served:

sudo mkdir -p /var/www/mta-sts/.well-known
sudoedit /var/www/mta-sts/.well-known/mta-sts.txt
caddy validate --config /etc/caddy/Caddyfile && sudo systemctl reload caddy
curl -sS https://mta-sts.example.com/.well-known/mta-sts.txt

Bump the id in the TXT record every time you change the policy file; senders cache the policy for max_age seconds and only re-fetch when the id changes.

Before moving from mode: testing to mode: enforce, verify every MX actually presents a valid, matching certificate — if your mail is hosted by a provider, you’re depending on their certificates:

for mx in $(dig +short MX example.com | awk '{print $2}' | sed 's/\.$//'); do
  echo "== $mx"
  # Needs outbound TCP/25; many cloud providers block it by default
  openssl s_client -starttls smtp -connect "$mx:25" -servername "$mx" \
    -verify_hostname "$mx" -verify_return_error -brief </dev/null 2>&1 \
    | grep -E 'Verification|Verified peername|verify error|Protocol|Peer certificate'
done

The TLS-RPT record routes daily reports about TLS failures when others deliver to you into the same mailbox, and parsedmarc lands them in the smtp_tls_* tables queried earlier. Run in testing mode until those reports are clean for a couple of weeks, then enforce.

Lessons Learned

p=none without someone reading the reports is theater. If nobody owns the report pipeline, you don’t have a DMARC deployment, you have a DNS record.

The DNS record is the easy part; the sender inventory is the project. Budget your time accordingly. Most of it goes to chasing down business units and SaaS admin panels, not to DNS.

DKIM is the leg that survives the real world. Every legitimate sender needs aligned DKIM. SPF-only passes are forwarding failures waiting for p=reject to expose them.

Relaxed alignment by default. Strict alignment is a precise tool for a closed environment, not a hardening checkbox.

Lock down non-sending domains immediately. p=reject on a parked domain has zero blast radius and removes an attack surface today.

Give reports their own mailbox and their own pipeline. parsedmarc, PostgreSQL and Grafana are enough, and the two queries that matter most — who’s failing, and who’s passing on SPF alone, fit on one screen.

Getting to p=reject isn’t a single change. It’s a sequence of measurable phases, and the data to move through them has been landing in your rua mailbox the whole time.

References


Written by Pablo Huertas - Cybersecurity Specialist & Technical Writer
19 years in enterprise security operations and incident response management.


Share this post:

Next Post
Green dashboard, broken control: what compliance automation doesn't fix