Walk me through how you would diagnose and resolve a production AIX server that is suddenly showing high CPU usage and causing application slowdowns.
A junior AIX administrator must be able to triage performance issues quickly and methodically. This question evaluates your troubleshooting approach, familiarity with AIX tools and commands, and ability to communicate actions under pressure — all critical when supporting Canadian enterprise environments (e.g., banks, telcos) where uptime is essential.
How to answer
- Begin with a clear, ordered approach (e.g., verify alert, gather facts, isolate cause, remediate, validate).
- Mention specific AIX commands and tools you would use (for example: topas, vmstat, iostat, ps -ef, errpt, nmon) and why.
- Explain how you'd distinguish CPU-bound processes vs. stuck kernel/thread issues vs. I/O wait or NUMA contention.
- Describe immediate mitigation steps (e.g., identify and throttle or restart offending processes, adjust priority with renice or chsysres/chdev if appropriate), and note safety checks before killing processes in production.
- Include steps for investigating root cause (application logs, recent deployments/changes, cron jobs, backups, kernel errors via errpt, patch level) and consult with application owners.
- Explain how you'd communicate status to stakeholders (clear, timely updates) and document actions for post-incident review.
- Finish with how you'd prevent recurrence (monitoring thresholds, runbook updates, capacity planning, patching or tuning suggestions).
What not to say
- Saying you'd immediately reboot the server without attempting investigation or graceful mitigation.
- Listing only generic commands without explaining what you would look for in their output.
- Claiming you'd 'kill all unknown processes' without caution or communication with app owners.
- Ignoring I/O or memory as possible contributors and focusing solely on CPU.
Sample answer
“First, I'd confirm the alert and note affected services and users. On the AIX host I'd run topas or nmon to view CPU, memory and I/O in real time, then ps -ef to find high-CPU processes and vmstat/iostat to check for I/O wait. If a specific process (e.g., a Java app) is consuming CPU after a recent deploy, I'd check the app logs and coordinate with the dev team before restarting the process. If I must act immediately to restore service, I'd gracefully stop the offending process or lower its priority with renice, and monitor impact. I'd also review errpt for kernel errors and check cron/jobs or backup windows. After service is stable I'd document findings, open a post-incident ticket, and suggest monitoring thresholds and a runbook update to prevent recurrence.”
Ready to rehearse this answer out loud?
Practice this question