Fix High CPU Usage on Linux: Step-by-Step Troubleshooting Guide

Updated 2026-08-14

Linux terminal command line
Photo: The GNOME Project / Wikimedia Commons (GPL)

High CPU doesn’t always mean a runaway process. First check load average against core count, then identify the culprit with top or htop, distinguish user from kernel CPU, and apply a targeted fix (kill process, renice, tune config). Most cases resolve in 10 minutes.

Understanding CPU Metrics

The Fundamental Difference: Load Average vs CPU Utilization

Two critical metrics often confused:

Load average (from uptime)

$ uptime
 10:45:23 up 45 days,  2:11, 2 users, load average: 2.50, 3.10, 2.85

On a 4-core system, a load average of 4.5 means tasks are queued and the system is oversubscribed. On an 8-core system, the same 4.5 is comfortable headroom.

CPU utilization (from top, htop, mpstat)

# Output from mpstat 1 1:
Linux 5.19.0 (prod-server)     08/14/2026

10:45:52 AM CPU    %usr   %nice   %sys  %iowait  %irq  %soft  %steal %idle
10:45:53 AM all   45.25   0.00   12.50   15.00   0.00   2.00   0.00  25.25
MetricMeaningAction if High
%usrUser-space app CPUProfile app, optimize code
%sysKernel CPU (syscalls, I/O)Check I/O bottleneck, tune config
%iowaitWaiting for disk/network I/OAdd faster storage, check I/O bandwidth
%irq / %softInterrupt handlingCheck for NIC/driver issues
%stealVM: host taking CPU awayMigrate to less-contended host
%idleAvailable capacityPlenty of headroom

Step 1: Get the Full Picture (5 Commands)

Run these in order to diagnose scope:

# 1. Check overall load and core count
uptime
nproc  # Number of CPU cores

# Example output:
#  09:15:33 up 12 days, 1:23, 3 users, load average: 8.50, 7.20, 6.10
#  8
# Interpretation: Load 8.5 on 8 cores = 106% utilization. System is overloaded.

# 2. Real-time CPU breakdown (five 1-second samples)
mpstat 1 5
# Tells you if it's user code, kernel, I/O wait, or steal time

# 3. Per-CPU breakdown (useful for single-threaded issues)
mpstat -P ALL 1 3

# 4. Process-level view
top -b -n 1 -o %CPU | head -20
# Or interactive:
top

# 5. Memory + process state
ps aux --sort=-%cpu | head -20

Real Example: Runaway Python Process

$ uptime
 09:15:33 up 12 days,  1:23, 3 users, load average: 8.50, 7.20, 6.10

$ nproc
8

$ mpstat 1 1
Linux 5.19.0 (web-server)     08/14/2026

09:15:35 AM CPU    %usr   %nice   %sys  %iowait  %irq  %soft  %steal %idle
09:15:36 AM all   87.50   0.00   10.00   0.00   0.00   0.00   0.00   2.50

$ ps aux --sort=-%cpu | head -5
USER       PID  %CPU %MEM    VSZ   RSS COMMAND
ubuntu   15678  95.2  2.1 1287956 172084 python3 /opt/ml/train.py
root        10   0.5  0.0      0     0 ksoftirqd/0

The Python job is the main culprit here, consuming nearly all available CPU. The smaller kernel-CPU share is worth a second look too, since it’s higher than idle but not the primary problem.

Step 2: Identify the Culprit (Process-Level Analysis)

Method 1: htop (Interactive, Best for Real-Time)

sudo apt-get install -y htop
htop

Navigate with arrow keys:

Method 2: ps (Batch, Best for Scripts)

# Top 10 CPU consumers with full command line
ps aux --sort=-%cpu | head -11

# Find all processes from a user
ps aux | grep ubuntu

# Find all processes in a cgroup (container)
ps -e --format pid,cgroup,comm | grep docker

# Show process tree (see parent-child relationships)
pstree -p | grep python

Method 3: top (Classic, No Install Required)

top

# Commands inside top:
# P         - Sort by %CPU
# M         - Sort by %MEM
# T         - Sort by TIME
# k         - Kill process (enter PID)
# r         - Renice process
# 1         - Show per-CPU breakdown

Method 4: Trace Syscalls (Deep Dive)

If top shows high %sys CPU, the process is making heavy syscalls:

# Trace system calls for PID 15678
strace -p 15678 -c

# After 5 seconds (Ctrl+C), you see:
# % time   seconds   usecs/call  calls    errors syscall
# ------ ---------- ----------- ----- -------- -----
# 45.20   3.214789       5       642341        read
# 32.10   2.285612       4       571892        write
# 18.50   1.316214       8       164521   EAGAIN epoll_wait
#  4.20   0.299124      11        27143        open

# Result: Massive read/write syscalls. Problem is likely I/O-bound, not CPU-bound.

Step 3: Categorize the Problem Type

Type A: CPU-Bound (App is Computing)

Causes: an inefficient algorithm, an infinite loop, a missing break condition, or an unoptimized regex.

Fixes:

# 1. Profile with perf (Linux)
sudo perf record -p 15678 -g -- sleep 10
sudo perf report
# Shows which functions burn CPU

# 2. Check for infinite loops
strace -p 15678 2>&1 | tail -50 | sort | uniq -c | sort -rn
# If you see hundreds of identical syscalls, it's spinning

# 3. Kill and restart with nice value
kill -9 15678
nice -n 19 python3 /opt/ml/train.py
# Lowest priority: won't starve other processes

Type B: I/O-Bound (Waiting for Disk/Network)

Causes: a slow disk, network congestion, NFS latency, or a database under heavy load.

Fixes:

# 1. Check disk I/O rate
iostat -x 1 3
# Look for high r/s, w/s, %util

# 2. Identify which file is slow
lsof -p 15678 | grep REG
# Check if it's /dev/sda (slow spinning) or /dev/nvme0n1 (fast SSD)

# 3. Check network traffic
nethogs -p

# 4. Move data to faster storage or add cache
# Example: Migrate database to SSD, add Redis layer

Type C: Kernel/Interrupt CPU

Causes: a faulty NIC driver, too many syscalls, or excessive context switching.

Fixes:

# 1. Check for interrupt storms
cat /proc/interrupts | sort -k2 -rn | head -5
# Watch for ballooning numbers on repeated runs

# 2. Check context switches
vmstat 1 3
# High `cs` column = thrashing. Too many processes?

# 3. Update kernel/drivers
sudo apt-get update && sudo apt-get install linux-image-generic
sudo reboot

# 4. Disable CPU features if urgent
echo 1 > /sys/kernel/debug/tracing/traceoff  # Disable tracing if it's a debug kernel

Type D: Container Throttling (if in Docker/K8s)

Fixes:

# Check if container is throttled
docker stats container_name
# Look for % CPU column

# Increase CPU limit
docker update --cpus 2 container_name
# Or in Kubernetes: edit Deployment spec.containers.resources.limits.cpu

Quick Fixes for Common Culprits

Fix 1: Kill a Runaway Process

# Find it
pgrep -f "python.*train.py"
# Output: 15678

# Kill it gracefully (SIGTERM)
kill 15678

# If it doesn't die in 10 seconds, force kill
kill -9 15678

# Auto-restart with systemd
sudo systemctl restart training-service

Fix 2: Renice a Legitimate But Hungry Process

Don’t kill it; just deprioritize:

# Current priority
ps -p 15678 -o pid,cmd,ni
# PID COMMAND                         NI
#15678 python3 /opt/ml/train.py       0

# Lower priority (higher number = less CPU time)
renice -n 15 -p 15678
# Now it runs only if nothing else needs CPU

# Verify
ps -p 15678 -o pid,cmd,ni
# PID COMMAND                         NI
#15678 python3 /opt/ml/train.py       15

Fix 3: Tune System Limits

If you have too many processes fighting for CPU:

# Check context switch rate
vmstat 1 3 | awk '{print "Context switches: " $12}'

# Reduce swappiness (prefer RAM to disk swap)
echo "vm.swappiness=10" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p

# Increase /proc/sys/kernel/sched_migration_cost_ns
# (Reduces unnecessary process migration between cores)
echo 5000000 | sudo tee /proc/sys/kernel/sched_migration_cost_ns

Fix 4: Check for Memory Pressure (Indirect CPU Rise)

High memory pressure forces swapping, which tanks CPU:

# Check swap usage
free -h
#              total        used        free
# Swap:        15Gi       12Gi        3.0Gi  <-- PROBLEM: 80% swap used

# Check page faults
vmstat 1 3 | grep "maj\|min"
# High minor faults (pflt) = page swapping

# Quick fix: Add more RAM or kill memory hogs

Preventive Monitoring

Set Up Alerts

Using a systemd timer (no external dependencies):

Create /etc/systemd/system/cpu-alert.service:

[Unit]
Description=CPU Usage Monitor
After=network.target

[Service]
Type=simple
ExecStart=/usr/local/bin/cpu-monitor.sh
Restart=always
RestartSec=10

[Install]
WantedBy=multi-user.target

Create /usr/local/bin/cpu-monitor.sh:

#!/bin/bash
THRESHOLD=80
LOAD=$(uptime | awk -F'load average:' '{print $2}' | awk '{print $1}' | sed 's/,//')
CORES=$(nproc)
UTILIZATION=$(echo "$LOAD * 100 / $CORES" | bc)

if (( $(echo "$UTILIZATION > $THRESHOLD" | bc -l) )); then
  echo "ALERT: CPU utilization ${UTILIZATION}% exceeds ${THRESHOLD}%" | mail -s "CPU Alert" ops@company.com
fi

Using Prometheus and Alertmanager, for a production setup:

# prometheus.yml alert rule
groups:
- name: cpu
  rules:
  - alert: HighCPUUsage
    expr: node_cpu_usage_percent > 80
    for: 5m
    annotations:
      summary: "CPU > 80% on {{ $labels.instance }}"

Baseline Metrics to Track

# Weekly check: Is CPU trend rising?
uptime
# Compare to last week

# Monthly: Are processes memory-leaking?
ps aux | grep YOUR_APP | grep -v grep

# Quarterly: Upgrade kernel?
uname -r
apt-cache policy linux-image-generic | head -5

Troubleshooting Checklist

Further reading

For more information about load average, see Load (computing) on Wikipedia. For more information about kernel-level CPU tuning, see the kernel’s own CPU management admin guide. For more information about the perf profiler used above, see the perf wiki.