Fix High CPU Usage on Linux: Step-by-Step Troubleshooting Guide
Updated 2026-08-14

High CPU doesn’t always mean a runaway process. First check load average against core count, then identify the culprit with top or htop, distinguish user from kernel CPU, and apply a targeted fix (kill process, renice, tune config). Most cases resolve in 10 minutes.
Understanding CPU Metrics
The Fundamental Difference: Load Average vs CPU Utilization
Two critical metrics often confused:
Load average (from uptime)
$ uptime
10:45:23 up 45 days, 2:11, 2 users, load average: 2.50, 3.10, 2.85
- 2.50 = average runnable tasks in last 1 minute
- 3.10 = average runnable tasks in last 5 minutes
- 2.85 = average runnable tasks in last 15 minutes
On a 4-core system, a load average of 4.5 means tasks are queued and the system is oversubscribed. On an 8-core system, the same 4.5 is comfortable headroom.
CPU utilization (from top, htop, mpstat)
# Output from mpstat 1 1:
Linux 5.19.0 (prod-server) 08/14/2026
10:45:52 AM CPU %usr %nice %sys %iowait %irq %soft %steal %idle
10:45:53 AM all 45.25 0.00 12.50 15.00 0.00 2.00 0.00 25.25
| Metric | Meaning | Action if High |
|---|---|---|
%usr | User-space app CPU | Profile app, optimize code |
%sys | Kernel CPU (syscalls, I/O) | Check I/O bottleneck, tune config |
%iowait | Waiting for disk/network I/O | Add faster storage, check I/O bandwidth |
%irq / %soft | Interrupt handling | Check for NIC/driver issues |
%steal | VM: host taking CPU away | Migrate to less-contended host |
%idle | Available capacity | Plenty of headroom |
Step 1: Get the Full Picture (5 Commands)
Run these in order to diagnose scope:
# 1. Check overall load and core count
uptime
nproc # Number of CPU cores
# Example output:
# 09:15:33 up 12 days, 1:23, 3 users, load average: 8.50, 7.20, 6.10
# 8
# Interpretation: Load 8.5 on 8 cores = 106% utilization. System is overloaded.
# 2. Real-time CPU breakdown (five 1-second samples)
mpstat 1 5
# Tells you if it's user code, kernel, I/O wait, or steal time
# 3. Per-CPU breakdown (useful for single-threaded issues)
mpstat -P ALL 1 3
# 4. Process-level view
top -b -n 1 -o %CPU | head -20
# Or interactive:
top
# 5. Memory + process state
ps aux --sort=-%cpu | head -20
Real Example: Runaway Python Process
$ uptime
09:15:33 up 12 days, 1:23, 3 users, load average: 8.50, 7.20, 6.10
$ nproc
8
$ mpstat 1 1
Linux 5.19.0 (web-server) 08/14/2026
09:15:35 AM CPU %usr %nice %sys %iowait %irq %soft %steal %idle
09:15:36 AM all 87.50 0.00 10.00 0.00 0.00 0.00 0.00 2.50
$ ps aux --sort=-%cpu | head -5
USER PID %CPU %MEM VSZ RSS COMMAND
ubuntu 15678 95.2 2.1 1287956 172084 python3 /opt/ml/train.py
root 10 0.5 0.0 0 0 ksoftirqd/0
The Python job is the main culprit here, consuming nearly all available CPU. The smaller kernel-CPU share is worth a second look too, since it’s higher than idle but not the primary problem.
Step 2: Identify the Culprit (Process-Level Analysis)
Method 1: htop (Interactive, Best for Real-Time)
sudo apt-get install -y htop
htop
Navigate with arrow keys:
F6(Sort): Click%CPUcolumn- View entire command with
<>arrows - Press
Sto strace the process kto kill,rto renice
Method 2: ps (Batch, Best for Scripts)
# Top 10 CPU consumers with full command line
ps aux --sort=-%cpu | head -11
# Find all processes from a user
ps aux | grep ubuntu
# Find all processes in a cgroup (container)
ps -e --format pid,cgroup,comm | grep docker
# Show process tree (see parent-child relationships)
pstree -p | grep python
Method 3: top (Classic, No Install Required)
top
# Commands inside top:
# P - Sort by %CPU
# M - Sort by %MEM
# T - Sort by TIME
# k - Kill process (enter PID)
# r - Renice process
# 1 - Show per-CPU breakdown
Method 4: Trace Syscalls (Deep Dive)
If top shows high %sys CPU, the process is making heavy syscalls:
# Trace system calls for PID 15678
strace -p 15678 -c
# After 5 seconds (Ctrl+C), you see:
# % time seconds usecs/call calls errors syscall
# ------ ---------- ----------- ----- -------- -----
# 45.20 3.214789 5 642341 read
# 32.10 2.285612 4 571892 write
# 18.50 1.316214 8 164521 EAGAIN epoll_wait
# 4.20 0.299124 11 27143 open
# Result: Massive read/write syscalls. Problem is likely I/O-bound, not CPU-bound.
Step 3: Categorize the Problem Type
Type A: CPU-Bound (App is Computing)
%usris high (>70%)%iowaitis low (<5%)%sysis normal (<15%)
Causes: an inefficient algorithm, an infinite loop, a missing break condition, or an unoptimized regex.
Fixes:
# 1. Profile with perf (Linux)
sudo perf record -p 15678 -g -- sleep 10
sudo perf report
# Shows which functions burn CPU
# 2. Check for infinite loops
strace -p 15678 2>&1 | tail -50 | sort | uniq -c | sort -rn
# If you see hundreds of identical syscalls, it's spinning
# 3. Kill and restart with nice value
kill -9 15678
nice -n 19 python3 /opt/ml/train.py
# Lowest priority: won't starve other processes
Type B: I/O-Bound (Waiting for Disk/Network)
%iowaitis high (>20%)%usris moderate (<50%)straceshowsread,write,pread64,sendto
Causes: a slow disk, network congestion, NFS latency, or a database under heavy load.
Fixes:
# 1. Check disk I/O rate
iostat -x 1 3
# Look for high r/s, w/s, %util
# 2. Identify which file is slow
lsof -p 15678 | grep REG
# Check if it's /dev/sda (slow spinning) or /dev/nvme0n1 (fast SSD)
# 3. Check network traffic
nethogs -p
# 4. Move data to faster storage or add cache
# Example: Migrate database to SSD, add Redis layer
Type C: Kernel/Interrupt CPU
%sysis high (>25%)%irqor%softis high (>5%)topshows ksoftirqd or kworker processes
Causes: a faulty NIC driver, too many syscalls, or excessive context switching.
Fixes:
# 1. Check for interrupt storms
cat /proc/interrupts | sort -k2 -rn | head -5
# Watch for ballooning numbers on repeated runs
# 2. Check context switches
vmstat 1 3
# High `cs` column = thrashing. Too many processes?
# 3. Update kernel/drivers
sudo apt-get update && sudo apt-get install linux-image-generic
sudo reboot
# 4. Disable CPU features if urgent
echo 1 > /sys/kernel/debug/tracing/traceoff # Disable tracing if it's a debug kernel
Type D: Container Throttling (if in Docker/K8s)
- Load is high but
topshows low CPU - Container has
cpuacct.cpu_throttled_count> 0
Fixes:
# Check if container is throttled
docker stats container_name
# Look for % CPU column
# Increase CPU limit
docker update --cpus 2 container_name
# Or in Kubernetes: edit Deployment spec.containers.resources.limits.cpu
Quick Fixes for Common Culprits
Fix 1: Kill a Runaway Process
# Find it
pgrep -f "python.*train.py"
# Output: 15678
# Kill it gracefully (SIGTERM)
kill 15678
# If it doesn't die in 10 seconds, force kill
kill -9 15678
# Auto-restart with systemd
sudo systemctl restart training-service
Fix 2: Renice a Legitimate But Hungry Process
Don’t kill it; just deprioritize:
# Current priority
ps -p 15678 -o pid,cmd,ni
# PID COMMAND NI
#15678 python3 /opt/ml/train.py 0
# Lower priority (higher number = less CPU time)
renice -n 15 -p 15678
# Now it runs only if nothing else needs CPU
# Verify
ps -p 15678 -o pid,cmd,ni
# PID COMMAND NI
#15678 python3 /opt/ml/train.py 15
Fix 3: Tune System Limits
If you have too many processes fighting for CPU:
# Check context switch rate
vmstat 1 3 | awk '{print "Context switches: " $12}'
# Reduce swappiness (prefer RAM to disk swap)
echo "vm.swappiness=10" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p
# Increase /proc/sys/kernel/sched_migration_cost_ns
# (Reduces unnecessary process migration between cores)
echo 5000000 | sudo tee /proc/sys/kernel/sched_migration_cost_ns
Fix 4: Check for Memory Pressure (Indirect CPU Rise)
High memory pressure forces swapping, which tanks CPU:
# Check swap usage
free -h
# total used free
# Swap: 15Gi 12Gi 3.0Gi <-- PROBLEM: 80% swap used
# Check page faults
vmstat 1 3 | grep "maj\|min"
# High minor faults (pflt) = page swapping
# Quick fix: Add more RAM or kill memory hogs
Preventive Monitoring
Set Up Alerts
Using a systemd timer (no external dependencies):
Create /etc/systemd/system/cpu-alert.service:
[Unit]
Description=CPU Usage Monitor
After=network.target
[Service]
Type=simple
ExecStart=/usr/local/bin/cpu-monitor.sh
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.target
Create /usr/local/bin/cpu-monitor.sh:
#!/bin/bash
THRESHOLD=80
LOAD=$(uptime | awk -F'load average:' '{print $2}' | awk '{print $1}' | sed 's/,//')
CORES=$(nproc)
UTILIZATION=$(echo "$LOAD * 100 / $CORES" | bc)
if (( $(echo "$UTILIZATION > $THRESHOLD" | bc -l) )); then
echo "ALERT: CPU utilization ${UTILIZATION}% exceeds ${THRESHOLD}%" | mail -s "CPU Alert" ops@company.com
fi
Using Prometheus and Alertmanager, for a production setup:
# prometheus.yml alert rule
groups:
- name: cpu
rules:
- alert: HighCPUUsage
expr: node_cpu_usage_percent > 80
for: 5m
annotations:
summary: "CPU > 80% on {{ $labels.instance }}"
Baseline Metrics to Track
# Weekly check: Is CPU trend rising?
uptime
# Compare to last week
# Monthly: Are processes memory-leaking?
ps aux | grep YOUR_APP | grep -v grep
# Quarterly: Upgrade kernel?
uname -r
apt-cache policy linux-image-generic | head -5
Troubleshooting Checklist
- Load average vs core count:
uptime; nproc - CPU breakdown (user/sys/iowait):
mpstat 1 5 - Top CPU consumer:
ps aux --sort=-%cpu | head - System call profile:
strace -p PID -c - Disk I/O status:
iostat -x 1 3 - Memory swap pressure:
free -h; vmstat 1 3 - Network traffic:
nethogs -p - Interrupt rate:
cat /proc/interrupts - Fix applied: Killed, reniced, or tuned?
- Monitoring in place: Alert rule created?
Further reading
For more information about load average, see Load (computing)
on Wikipedia. For more information about kernel-level CPU tuning, see the
kernel’s own CPU management admin guide.
For more information about the perf profiler used above, see the
perf wiki.
☕ Coffee Corner
Compiling, deploying, or waiting on a render? Here's what to brew while you wait.
🏠 Smart Home Picks
Same hobbyist care applied to your network and your front door.




