Critical warning decision tool

NVMe SMART Warning Triage

Turn a smartd Critical Warning email into an ordered triage: decode the warning bit, separate a live thermal state from a remembered smartd alert, and get the exact replace-now signals from your SMART counters.

Decision-first: only the fields that change what you should do. For a full report walk-through use the health explainer instead.

From scary email to a decision in five fields

A smartd email reading Device: /dev/nvme0, Critical Warning (0x02): Temperature looks like a failing drive and usually is not. The triage above turns the alert into one of three verdicts: cool the drive (live thermal state), reset the daemon (remembered state), or back up and plan the replacement (integrity signals). It asks only for the fields that change the verdict — the current warning byte, the two composite temperature counters, media errors, spare headroom and sensor agreement.

This tool is the decision layer on top of our NVMe Critical Warning 0x02 guide, which explains every field with real smartctl captures. If you want the whole report translated rather than a verdict, the SMART & NVMe Health Explainer does that locally in your browser.

The five warning bits, ranked by urgency

The Critical Warning byte packs five independent flags into eight bits. They are not equally dangerous, and the triage weights them differently:

Critical Warning bits per the NVMe SMART/Health log (log identifier 02h)
BitMeaningVerdict weight
0x01Available spare capacity has fallen below the thresholdEscalate — write headroom is gone
0x02Composite temperature above an over-temperature or below an under-temperature thresholdCool first — reversible live state
0x04NVM subsystem reliability degradedEscalate — internal wear or fault
0x08All media placed in read-only modeEscalate — backup immediately
0x10Volatile memory backup device failedEscalate — power-loss write risk

Only 0x02 clears itself when conditions improve. Every other bit, or any combination containing 0x04/0x08/0x10, routes the triage to the replace path regardless of temperature.

Live state vs remembered state: the smartd trap

The most confusing case is an alert that keeps coming back for a drive that is currently cool. smartd keeps its own state and repeats the warning every 24 hours, so a drive that overheated once can be emailed about indefinitely. In a documented Proxmox case, a Samsung SSD 980 1TB read 0x00 at 43 Celsius while Warning Composite Temperature Time read 748 minutes — the residue of a room cooling failure, not a live problem. Restarting the smartmontools service cleared the remembered alert.

The triage treats this pattern as its own verdict: current byte 0x00, alert repeating, counter above zero. The counters stay above zero on purpose — they are cumulative lifetime totals and the only record of an excursion nobody watched. Read them as history, not as a current state.

The healthy baseline to compare against

The triage anchors its verdicts to a measured never-throttled baseline from our test bed, kept in the public measurement repo:

EQi12 internal NVMe healthy baseline (source: EQi12 Measurement Data repo, storage/EQi12_InternalSSD_SMART_2026-07-11.txt, smartctl 7.5, CC BY 4.0)
FieldMeasured valueCondition
Critical Warning0x00Idle desktop workload
Temperature38 C (both sensors agree)Passive mini PC chassis, idle
Available Spare / Threshold100% / 10%New drive, 89 power-on hours
Warning / Critical Composite Temp Time0 / 0 minutesNo thermal excursion in drive life
Media and Data Integrity Errors0Read-only SMART check

Two sensors that agree within about two degrees is normal on this class of hardware. A sensor that diverges from its partner by more than a few degrees points at a mounting or contact problem — worth fixing, but not grounds for replacement on its own.

Commands the checklist relies on

smartctl -a /dev/nvme0           # full SMART/Health log — read every triage field here
smartctl -l error /dev/nvme0     # error information log — should stay at zero entries
systemctl restart smartmontools  # clears a remembered 0x02 state after cooling is fixed

On Windows the same drive is reachable as a physical device, for example smartctl -a -d nvme,0 /dev/pd0 with smartmontools installed. After any cooling change, re-run the triage with fresh values and record the date — the point is a before/after pair of counter readings, not a single snapshot.

The replace signals, in one list

  • Media and Data Integrity Errors above zero — any value is a backup-now signal.
  • Available Spare at or below its threshold (commonly 10%) — write endurance headroom is gone.
  • Error Information Log Entries climbing across consecutive reads.
  • Any 0x04, 0x08 or 0x10 bit set — these do not reverse.

A temperature warning with none of the above is a cooling project: reposition the machine, clear intake, re-check in a week, and let the composite counters confirm the excursion stopped growing. If a replacement does become necessary, size the migration with the backup retention planner first — a replace decision without a tested backup path is how one bad drive becomes two problems.

Frequently asked questions

Is NVMe Critical Warning 0x02 an emergency?

Usually not — it is a live, reversible state that clears when the drive cools. Escalate on 0x04/0x08/0x10, media errors above zero, or a spare at its threshold. The full guide walks the spec definition field by field.

Why does smartd keep emailing 0x02 when smartctl shows 0x00?

smartd remembers the warning and repeats it every 24 hours. A Samsung 980 1TB case showed 0x00 at 43 C with a 748-minute warning-time residue; restarting smartmontools cleared it.

Which counters prove a thermal excursion really happened?

Warning and Critical Composite Temperature Time — cumulative lifetime minutes above each threshold. The EQi12 healthy baseline reads zero in both at 38 C.

When do I replace instead of just cooling?

On integrity signals: media errors, spare at threshold, climbing error log entries, or bits 0x04/0x08/0x10. Temperature alone is never the replace trigger.