Debugging and Engineering Tools : From A Junior to Senior
How senior engineers find what is actually wrong
A beginner often thinks debugging means:
“The program crashed. I will add some print statements.”
A stronger engineer thinks:
“What evidence can I collect to narrow down the failure?”
A senior engineer goes one step further:
“At which layer is the system violating my mental model?”
That layer might be:
source code,
process state,
memory,
filesystem,
operating system,
network,
container,
GPU,
ROS graph,
device driver,
sensor,
electrical signal,
timing,
or physical hardware.
This is why debugging tools matter.
They let you observe layers of a system that are otherwise invisible.
A useful mental model is:
Hypothesis→Measurement→Evidence→Narrower Hypothesis\boxed{ \text{Hypothesis} \rightarrow \text{Measurement} \rightarrow \text{Evidence} \rightarrow \text{Narrower Hypothesis} }
For example:
“The API is slow.”
That statement alone tells us almost nothing.
Possible causes include:
application code is slow,
database is slow,
DNS resolution is slow,
network latency is high,
packets are retransmitted,
GPU inference is saturated,
CPU is throttled,
disk I/O is blocking,
lock contention exists,
another service is overloaded.
Different tools reveal different causes.
That is the essence of engineering diagnostics.
1. The Debugging Mindset
Before learning any tool, learn the method.
Suppose a robot suddenly stops moving.
A weak debugging process is:
restart ROS
restart Gazebo
reboot computer
change random parameters
Sometimes that works.
But you learn almost nothing.
A systematic process asks:
Is the process alive?
Is the node publishing commands?
Is
/cmd_velreceiving messages?Is the controller subscribed?
Is the motor driver receiving commands?
Is communication reaching the hardware?
Is electrical power available?
Are the motors physically capable of moving?
Each step eliminates an entire category of possibilities.
Good debugging is fundamentally search-space reduction.
If the original set of possible causes is:
C={c1,c2,…,c1000}C = \{c_1,c_2,\ldots,c_{1000}\}
a useful diagnostic test partitions it:
C→Cpossible⊂CC \rightarrow C_{possible} \subset C
A strong test removes many possibilities at once.
2. Git
Git is usually described as a version-control system.
That is correct, but incomplete.
Git is really a system for recording snapshots of software history and relationships between those snapshots.
You can think of your project as evolving through states:
S0→S1→S2→S3S_0 \rightarrow S_1 \rightarrow S_2 \rightarrow S_3
A Git commit records one of those states.
2.1 Why Git matters for debugging
Suppose your program worked yesterday but fails today.
Instead of manually examining 500 changed lines, you can ask:
What changed between the working state and the broken state?
git diff <working-commit> <broken-commit>
This converts debugging from:
search the whole system
into:
search the changed region.
That is enormously powerful.
2.2 The three important Git states
A practical model is:
Working tree→Staging area→Repository\text{Working tree} \rightarrow \text{Staging area} \rightarrow \text{Repository}
Working tree
Files currently on disk.
Staging area
Changes selected for the next commit.
Repository
Previously committed history.
Example:
git status
might show:
modified: backend/app.py
After:
git add backend/app.py
the change enters the staging area.
After:
git commit -m "Fix inference timeout"
it becomes part of permanent Git history.
2.3 Branches
A branch is essentially a movable reference to a commit.
Imagine:
A --- B --- C
^
main
Create a feature branch:
A --- B --- C
\
D --- E
feature
This lets development diverge safely.
2.4 git diff
One of the most important debugging tools:
git diff
Shows uncommitted changes.
git diff --staged
Shows staged changes.
git diff HEAD~1 HEAD
Shows what changed in the latest commit.
2.5 git log
git log --oneline --graph --decorate
gives a compact history.
Example:
* a31df12 Fix scoring threshold
* f18dd22 Add evidence resolver
* e2715c5 Initial binding implementation
You can now reason about software evolution.
2.6 git blame
git blame backend/scorer.py
shows which commit introduced each line.
Despite its unfortunate name, the purpose is not blaming people.
It answers:
When and why did this code appear?
2.7 git bisect
One of Git's most powerful debugging features.
Suppose commit:
A
works.
Commit:
Z
fails.
There may be 100 commits between them.
Instead of testing all 100, Git performs binary search.
git bisect start
git bisect bad Z
git bisect good A
Git checks a middle commit.
You test it:
git bisect good
or:
git bisect bad
After only approximately:
log2(100)≈7\log_2(100) \approx 7
tests, Git can identify the offending commit.
That is a beautiful application of binary search to debugging.
3. Linux
For AI, robotics, cloud systems, servers, containers, GPUs, and research infrastructure, Linux is not optional knowledge.
It is the environment in which much of serious computing actually runs.
The important thing is not memorizing commands.
It is understanding the operating-system model.
3.1 Everything revolves around processes
A running program is a process.
View processes:
ps aux
Interactive view:
top
or:
htop
A process has:
PID,
memory,
CPU state,
open files,
environment variables,
network sockets,
threads.
For example:
ps aux | grep python
may reveal multiple Python services when you expected one.
3.2 Signals
Processes receive signals.
Examples:
SIGTERM
SIGKILL
SIGINT
SIGSEGV
When you press:
Ctrl+C
the terminal normally sends:
SIGINT
A graceful shutdown typically uses:
kill <PID>
which usually sends SIGTERM.
Force termination:
kill -9 <PID>
sends SIGKILL.
SIGKILL cannot be handled by the application.
So it should not be your first debugging tool.
3.3 Files and descriptors
Linux represents many things through file descriptors.
Processes can have:
files,
sockets,
pipes,
devices
open simultaneously.
You can inspect them with:
lsof -p <PID>
This is useful when asking:
Which file is this process using?
or:
Which port does it have open?
3.4 /proc
Linux exposes enormous amounts of process and kernel information through:
/proc
For example:
cat /proc/<PID>/status
can show process memory and state.
cat /proc/cpuinfo
shows CPU information.
cat /proc/meminfo
shows memory statistics.
This reveals an important Linux philosophy:
system state should be inspectable.
4. SSH
SSH means Secure Shell.
It provides encrypted remote access.
Conceptually:
Your machine
|
encrypted connection
|
Remote server
Typical command:
ssh user@server
4.1 Password authentication
You enter a password.
Simple but inconvenient for automation.
4.2 Public-key authentication
You generate a key pair:
private key
public key
The private key stays with you.
The public key is placed on the server.
Authentication essentially proves:
I possess the private key corresponding to this public key.
Create a key:
ssh-keygen -t ed25519
4.3 SSH is more than remote shell
Port forwarding is extremely useful.
Suppose a remote service listens only on:
127.0.0.1:8000
on a server.
You can tunnel it:
ssh -L 8000:localhost:8000 user@server
Then accessing:
localhost:8000
on your laptop forwards through SSH to the remote service.
This is common for:
Jupyter,
dashboards,
database tools,
internal APIs.
5. The Terminal
The terminal is not merely a place to type commands.
It is a way of composing programs.
Unix tools are often designed around:
input
→ transformation
→ output
This lets programs be connected together.
5.1 Standard streams
A process commonly has:
stdin
stdout
stderr
Standard input:
stdin
Standard output:
stdout
Errors:
stderr
5.2 Redirection
Write output to a file:
python app.py > output.log
Errors only:
python app.py 2> error.log
Both:
python app.py > output.log 2>&1
5.3 Pipes
A pipe connects one command's output to another's input:
command1 | command2
Example:
ps aux | grep python
Meaning:
list processes
→ filter lines containing "python"
This composability is central to Unix engineering.
6. grep and ripgrep
Logs and source trees quickly become enormous.
Search tools let you ask precise questions.
6.1 grep
grep "ERROR" app.log
Find all matching lines.
Recursive search:
grep -R "timeout" .
Case-insensitive:
grep -Ri "timeout" .
Line numbers:
grep -Rn "timeout" .
6.2 ripgrep
ripgrep, usually invoked as:
rg
is generally faster and friendlier for source-code search.
Example:
rg "HYPERCLOVAX_URL"
Search only Python:
rg "timeout" -g "*.py"
Show surrounding context:
rg -C 3 "connection refused"
A senior engineer frequently uses search before opening files manually.
6.3 Search is part of debugging
Suppose logs contain:
request_id=913fa2
Search every log:
rg "913fa2" logs/
Now you can reconstruct one request's journey through multiple components.
This is the basic idea behind distributed tracing as well.
7. Build Systems
Compiling large software involves many dependencies.
Suppose:
main.cpp
depends on
planner.cpp
depends on
planner.hpp
A build system tracks these relationships.
Instead of recompiling everything every time, it determines what changed.
Conceptually:
source files+dependencies+compiler options→binary\text{source files} + \text{dependencies} + \text{compiler options} \rightarrow \text{binary}
Examples include:
Make,
Ninja,
Bazel,
Meson,
CMake-generated systems.
7.1 Why build systems matter
Imagine 10,000 C++ files.
Only one header changes.
A good build system determines which targets depend on that header and rebuilds only what is necessary.
This can reduce build time dramatically.
8. CMake
CMake is technically a build-system generator.
It typically generates build files for:
Ninja,
Make,
Visual Studio,
others.
A simple CMakeLists.txt:
cmake_minimum_required(VERSION 3.16)
project(robot_controller)
add_executable(robot
main.cpp
controller.cpp
)
Then:
cmake -S . -B build
cmake --build build
8.1 Libraries
add_library(controller
controller.cpp
)
Then:
target_link_libraries(robot PRIVATE controller)
This describes dependency relationships.
8.2 Include directories
target_include_directories(
controller
PUBLIC include
)
8.3 Why modern CMake prefers targets
Older CMake often configured global flags.
Modern CMake thinks in terms of targets:
target
├── include paths
├── compiler flags
├── dependencies
└── linked libraries
This prevents build settings from leaking unpredictably through the project.
9. Debuggers
A debugger allows you to stop a running program and inspect internal state.
Instead of guessing:
Maybe
xhas the wrong value.
you can stop execution and inspect:
x = ?
Core debugger capabilities include:
breakpoints,
stepping,
variable inspection,
call stacks,
memory inspection,
watchpoints.
10. GDB
GDB is the standard debugger for native Linux programs.
Compile with debug information:
gcc -g program.c -o program
Launch:
gdb ./program
10.1 Breakpoints
break main
or:
break controller.cpp:42
Run:
run
The program pauses at that location.
10.2 Stepping
Execute one source line:
next
Step inside a function:
step
Continue normally:
continue
10.3 Variable inspection
print velocity
Maybe:
$1 = -1.742e+38
Now you immediately know something is seriously wrong.
10.4 Call stacks
backtrace
Example:
#0 divide()
#1 compute_score()
#2 process_request()
#3 main()
This tells you exactly how execution reached the failure.
10.5 Core dumps
When a program crashes, Linux can save process state as a core dump.
Then you can inspect the crash afterward:
gdb ./program core
This is extremely useful for production crashes that are difficult to reproduce interactively.
11. Sanitizers
Some bugs corrupt memory without crashing immediately.
These are among the most dangerous C/C++ bugs.
Sanitizers instrument your program to detect them.
11.1 AddressSanitizer
Compile:
-fsanitize=address
Detects problems such as:
heap buffer overflow,
stack buffer overflow,
use-after-free,
double free.
Example:
int a[5];
a[7] = 10;
This writes outside the array.
Without a sanitizer, the program may appear to work.
With AddressSanitizer, you receive a detailed report.
11.2 UndefinedBehaviorSanitizer
-fsanitize=undefined
Detects undefined behavior such as:
invalid shifts,
some integer overflow,
misaligned access,
invalid casts.
11.3 ThreadSanitizer
-fsanitize=thread
Detects data races.
Suppose two threads modify:
counter++;
simultaneously without synchronization.
That operation is not necessarily atomic.
ThreadSanitizer can catch the race.
12. Profilers
Debuggers answer:
Why is the program incorrect?
Profilers answer:
Where is the program spending its resources?
A profiler can measure:
CPU time,
wall-clock time,
memory allocations,
call frequency,
cache behavior,
GPU activity.
12.1 Sampling profiler
A sampling profiler periodically asks:
What function is executing right now?
Suppose after 100,000 samples:
matrix_multiply 60%
tokenize 5%
network_wait 20%
other 15%
Now you know where optimization effort matters.
12.2 Flame graphs
Flame graphs visualize call stacks.
Wide regions consume more CPU time.
Conceptually:
main
├── inference
│ ├── attention
│ └── matmul
└── preprocessing
If matmul occupies most horizontal width, that is where most sampled time is spent.
12.3 Do not optimize before profiling
A common mistake:
This loop looks slow. I'll rewrite it.
But maybe it consumes only 1% of runtime.
Even making it infinitely fast gives at most approximately:
1%1\%
overall improvement.
Profiling prevents wasted optimization effort.
13. Valgrind-Style Tools
Valgrind is a dynamic-analysis framework.
Its famous tool, Memcheck, detects:
memory leaks,
invalid reads,
invalid writes,
use of uninitialized memory.
Run:
valgrind --leak-check=full ./program
Example output may reveal:
100 bytes definitely lost
meaning allocated memory was never freed.
13.1 Sanitizer versus Valgrind
They overlap but work differently.
Sanitizers:
compile-time instrumentation,
generally faster,
excellent integration during development.
Valgrind-style instrumentation:
often requires no special compilation,
can provide deep runtime checking,
typically much slower.
Modern C/C++ development often uses sanitizers heavily, while Valgrind remains useful for certain diagnostics.
14. Packet Capture and Wireshark
When applications communicate over networks, bugs may happen below the application layer.
Wireshark lets you inspect network packets.
Conceptually:
application
↓
TCP/UDP
↓
IP
↓
Ethernet/Wi-Fi
Wireshark lets you observe these layers directly.
14.1 What packet capture can answer
Questions like:
Did the request actually leave the machine?
Did the server reply?
Is TCP retransmitting packets?
Is DNS failing?
Is the connection being reset?
Are packets arriving out of order?
14.2 tcpdump
On Linux:
sudo tcpdump -i any port 8000
captures traffic involving port 8000.
Save packets:
sudo tcpdump -i eth0 -w capture.pcap
Then open the .pcap file in Wireshark.
14.3 TCP handshake
A normal TCP connection begins:
Client → SYN
Server → SYN-ACK
Client → ACK
If you see:
SYN
SYN
SYN
with no response, the server or network may be unreachable.
If you see:
SYN
RST
the destination may be actively rejecting the connection.
This lets you distinguish:
application bug
from:
network connectivity bug
15. ROS Diagnostics
ROS systems are distributed systems.
A robot may contain dozens of nodes exchanging hundreds of topics.
That creates enormous debugging complexity.
You need to inspect the graph.
15.1 Nodes
ros2 node list
Shows running nodes.
If your expected node is missing, debugging starts there.
15.2 Topics
ros2 topic list
Inspect messages:
ros2 topic echo /cmd_vel
Suppose the robot does not move.
If /cmd_vel contains correct velocity commands:
linear:
x: 0.2
angular:
z: 0.0
then the planner is likely functioning.
The failure lies farther downstream.
15.3 Topic information
ros2 topic info /cmd_vel
shows publishers and subscribers.
If:
Publisher count: 1
Subscription count: 0
then commands are being produced but nobody is listening.
15.4 Topic frequency
ros2 topic hz /scan
Maybe expected:
10 Hz
but actual:
1.2 Hz
That could explain localization instability.
15.5 Services
ros2 service list
Inspect:
ros2 service type /some_service
Call:
ros2 service call ...
15.6 Parameters
ros2 param list
Get one:
ros2 param get /controller_server max_vel_x
Many ROS bugs come from configuration rather than code.
15.7 TF debugging
Transforms are essential in robotics.
Typical frame chain:
map
↓
odom
↓
base_link
↓
laser
If:
map → odom
is missing, navigation can fail even though all nodes are alive.
Useful tools include:
ros2 run tf2_ros tf2_echo map base_link
and visual TF trees.
The important debugging principle is:
Check dataflow, not merely process existence.
16. Tracing
Logs tell you what individual components say.
Tracing tells you how work travels across the system.
Imagine:
API Gateway
→ Scoring Service
→ Model Service
→ Database
One request receives an ID:
request_id = abc123
Every subsystem propagates that identifier.
Now you can reconstruct:
abc123
Gateway 2 ms
Scoring 12 ms
Model 1840 ms
Database 4 ms
Immediately:
Model inference dominates latency.
16.1 Spans
Distributed tracing uses spans.
Example:
request
├── parse_input
├── retrieve_context
├── inference
│ ├── tokenize
│ ├── forward_pass
│ └── decode
└── serialize
Each span has timing information.
16.2 Why traces outperform logs for latency debugging
Logs may say:
Started request
Called model
Finished request
But tracing captures structured parent-child timing relationships.
This is invaluable for microservices and AI inference systems.
17. Log Analysis
Logs are the historical record of a running system.
But logging every possible thing is not good logging.
Useful logs answer questions.
17.1 Good log structure
Bad:
Something failed
Better:
ERROR inference_failed
request_id=abc123
model=hyperclovax
status=503
latency_ms=4021
retry=2
Structured logs make searching and aggregation much easier.
17.2 Log levels
Common levels:
TRACE
DEBUG
INFO
WARN
ERROR
FATAL
A rough interpretation:
TRACE
Extremely detailed execution information.
DEBUG
Useful developer diagnostics.
INFO
Normal significant events.
WARN
Something unusual happened, but the system continues.
ERROR
An operation failed.
FATAL
System cannot continue.
17.3 Correlation IDs
For distributed systems, attach a unique identifier:
request_id
trace_id
session_id
Then:
rg "abc123" logs/
reconstructs the lifecycle of that request.
17.4 Logs need context
Bad:
timeout
Good:
model_request_timeout
request_id=abc123
timeout_ms=30000
elapsed_ms=30021
endpoint=/v1/chat/completions
The goal is to record enough evidence to debug after the failure has already occurred.
18. GPU Profiling
AI workloads frequently move the bottleneck from CPU to GPU.
At that point ordinary CPU profiling is not enough.
You need to inspect:
kernel execution,
GPU utilization,
memory bandwidth,
memory transfers,
synchronization,
tensor-core usage.
18.1 GPU utilization
A common NVIDIA tool:
nvidia-smi
It shows things such as:
GPU utilization,
VRAM usage,
temperature,
processes.
Suppose:
GPU-Util: 15%
Memory: 22GB / 24GB
You are using lots of memory but little computation.
Possible causes:
CPU bottleneck,
small batches,
synchronization,
slow data loading,
inefficient kernels.
18.2 Kernel execution
A GPU program launches kernels.
Conceptually:
CPU
→ launch kernel
→ GPU executes
→ synchronization
If kernels are tiny and numerous, launch overhead can dominate.
Profilers can reveal timelines like:
kernel A: 20 µs
gap: 100 µs
kernel B: 18 µs
gap: 95 µs
The GPU spends more time waiting than computing.
18.3 Memory-bound versus compute-bound
Suppose a kernel performs little arithmetic but reads enormous memory.
Then performance is limited by:
memory bandwidth\text{memory bandwidth}
not:
compute throughput\text{compute throughput}
Another kernel may perform huge matrix multiplications and become compute-bound.
Knowing which regime you are in determines how optimization should proceed.
18.4 CPU-GPU transfers
Moving data between CPU and GPU is expensive.
A timeline might show:
CPU → GPU copy
GPU compute
GPU → CPU copy
GPU compute
CPU → GPU copy
Too many transfers can destroy performance.
18.5 Synchronization
Operations such as:
tensor.cpu()
or certain timing calls may force the CPU to wait until the GPU finishes.
This can serialize what should have been asynchronous work.
GPU profiling exposes these stalls.
19. Hardware Debugging
Software engineers often assume:
If the software says it happened, the hardware did it.
Embedded and robotics engineers know better.
A command may be sent correctly in software but never reach the physical device.
You may need to debug:
application
→ driver
→ communication bus
→ voltage signal
→ device
19.1 Basic checks
Always begin with simple physical questions:
Is power present?
Is ground connected?
Is voltage correct?
Are cables connected?
Is polarity correct?
Is the device overheating?
Is a connector intermittent?
These may sound trivial.
They are extremely common failure sources.
19.2 Multimeter
A multimeter measures:
voltage,
current,
resistance,
continuity.
For example:
Expected:
5.0V5.0V
Measured:
1.3V1.3V
You have immediately found a hardware-level problem.
No amount of software debugging will fix that.
20. Oscilloscope
An oscilloscope displays voltage over time.
Conceptually:
V(t)V(t)
Instead of asking:
Is the pin high or low?
you can observe:
How does the voltage evolve over microseconds or nanoseconds?
20.1 Digital pulse example
Suppose a microcontroller should produce:
HIGH
LOW
HIGH
LOW
at 1 kHz.
An oscilloscope might show:
period ≈ 1 ms
confirming the signal frequency.
20.2 Signal integrity
Real signals are not perfect squares.
You may observe:
ringing,
overshoot,
undershoot,
noise,
slow rise time,
jitter.
Example:
ideal:
____|‾‾‾‾|____
actual:
____/\/\/‾\____
At high speeds, these imperfections can cause communication errors.
20.3 Analog signals
For sensor outputs:
voltage
↑
| /\ /\
| / \____/ \
|___/______________→ time
An oscilloscope lets you inspect actual sensor behavior directly.
21. Logic Analyzer
A logic analyzer is designed for digital signals.
Instead of displaying analog voltage shape in detail, it interprets lines primarily as:
0
1
over time.
This makes it excellent for protocols.
21.1 UART
A logic analyzer can decode serial data such as:
0x48 0x65 0x6C 0x6C 0x6F
which corresponds to:
Hello
Now you can verify exactly what bytes crossed the wire.
21.2 I²C
You can inspect:
START
address
ACK
data
ACK
STOP
If the device fails to acknowledge:
NACK
you may have:
wrong device address,
wiring problem,
unpowered device,
timing issue.
21.3 SPI
You can inspect:
MOSI
MISO
CLK
CS
and verify whether transmitted bits match expectations.
21.4 Oscilloscope versus logic analyzer
Use an oscilloscope when you care about:
electrical waveform
Use a logic analyzer when you care about:
digital protocol behavior
Often engineers use both.
For example:
Logic analyzer says:
Data bits are wrong.
Oscilloscope says:
The clock edges are distorted due to signal-integrity problems.
22. Performance Benchmarking
Benchmarking means measuring system performance reproducibly.
This sounds simple.
It is surprisingly easy to do badly.
22.1 Latency
Latency measures how long one operation takes.
For request ii:
Li=tfinish−tstartL_i = t_{finish}-t_{start}
Do not report only average latency.
Real systems often care about percentiles:
p50
p90
p95
p99
Example:
p50 = 120 ms
p95 = 260 ms
p99 = 1800 ms
The average might look acceptable while 1% of users experience severe delays.
22.2 Throughput
Throughput measures work per unit time.
Examples:
requests/second\text{requests/second}tokens/second\text{tokens/second}images/second\text{images/second}samples/second\text{samples/second}
Latency and throughput are related but not identical.
A system might increase batching:
batch = 1
→ 20 ms latency
→ 50 requests/s
versus:
batch = 32
→ 100 ms latency
→ 800 requests/s
Higher throughput, worse individual latency.
22.3 Warm-up
The first operation may be slower because of:
library initialization,
CUDA context creation,
JIT compilation,
caching,
disk page faults.
Therefore benchmark:
warm-up
→ repeated measurements
→ statistics
not merely one run.
22.4 Control the environment
A proper benchmark should specify:
hardware,
CPU model,
GPU model,
software version,
batch size,
precision,
number of threads,
input dimensions,
power settings,
compiler flags.
Otherwise results may not be reproducible.
22.5 Compare apples to apples
Suppose:
Model A:
FP16
batch 32
A100 GPU
Model B:
FP32
batch 1
RTX 4060
Comparing their inference speed tells you almost nothing about the models themselves.
Benchmarks must control variables.
23. How the Tools Fit Together
Consider this production incident:
Your AI API suddenly takes 12 seconds per request instead of 2 seconds.
A mature investigation may proceed like this.
Step 1 — Logs
You discover:
model_inference_ms=10750
Most latency occurs in inference.
Step 2 — Linux process inspection
top
nvidia-smi
CPU is at 100%.
GPU utilization is only 18%.
Interesting.
Step 3 — CPU profiling
Profiler shows:
tokenization 68%
GPU compute 19%
other 13%
Now the bottleneck is obvious.
Step 4 — Git
git diff HEAD~5 HEAD
reveals that a recent commit changed the tokenizer implementation.
Step 5 — Git bisect
You identify the exact commit that introduced the regression.
Step 6 — Benchmark
Before fix:
p50 = 11.8 s
After fix:
p50 = 2.1 s
You have not merely "fixed something."
You have:
measured the failure,
localized the bottleneck,
identified the cause,
repaired it,
validated the improvement.
That is professional debugging.
24. A Robotics Example
Suppose a robot is not responding to navigation commands.
A systematic chain might be:
Is Nav2 running?
ros2 node list
Yes.
Are velocity commands produced?
ros2 topic echo /cmd_vel
Yes.
Is the motor controller subscribed?
ros2 topic info /cmd_vel
Yes.
Are command rates correct?
ros2 topic hz /cmd_vel
Yes.
Is the serial driver transmitting?
Check driver logs.
Yes.
Are bytes physically present?
Use logic analyzer.
No.
Now the problem is between:
driver
and
physical UART output
Maybe the wrong device is opened.
Maybe GPIO configuration is wrong.
Maybe the UART peripheral is disabled.
The debugging space has been reduced from an entire autonomous robot to one tiny subsystem.
25. Debugging AI Models
Not every AI bug is a software crash.
Suppose model quality suddenly decreases.
The debugging sequence may be:
Git
What training code changed?
Dataset checks
Did preprocessing change?
Logs
Did loss behave differently?
Metrics
Which class degraded?
GPU profiling
Was training unexpectedly slower?
Benchmarking
Did throughput change?
Experiment tracking
Did learning rate or batch size change?
Statistical analysis
Is the degradation larger than random-run variance?
Model inspection
Did gradients explode?
The same debugging principles apply:
observe→measure→isolate→test\text{observe} \rightarrow \text{measure} \rightarrow \text{isolate} \rightarrow \text{test}
26. Common Failure Patterns and the Tool to Reach For
“The program crashes.”
Think:
GDB,
core dump,
sanitizers.
“Memory keeps increasing.”
Think:
Valgrind-style memory tools,
heap profiler,
sanitizers.
“The application is slow.”
Think:
profiler,
tracing,
benchmarks.
“The server cannot connect.”
Think:
curl,ss,tcpdump,Wireshark,
DNS inspection.
“It worked yesterday.”
Think:
Git diff,
Git log,
Git bisect.
“The ROS robot is doing nothing.”
Think:
node graph,
topics,
TF,
services,
parameters,
diagnostics.
“GPU memory is full.”
Think:
nvidia-smi,GPU memory profiler,
tensor-lifetime analysis.
“GPU utilization is low.”
Think:
GPU profiler,
data-loading profiler,
CPU-GPU synchronization,
batching.
“The embedded device ignores commands.”
Think:
logs,
serial capture,
logic analyzer,
oscilloscope.
“Performance numbers are inconsistent.”
Think:
benchmarking methodology,
warm-up,
hardware state,
statistical variance.
27. Observability Layers
A useful senior-level mental model is to think of debugging as moving through layers.
┌──────────────────────────────┐
│ Application behavior │
│ logs / exceptions / metrics │
├──────────────────────────────┤
│ Runtime │
│ debugger / profiler / trace │
├──────────────────────────────┤
│ Operating system │
│ processes / memory / files │
├──────────────────────────────┤
│ Network │
│ sockets / packets / DNS │
├──────────────────────────────┤
│ Accelerator │
│ GPU kernels / memory / sync │
├──────────────────────────────┤
│ Robotics middleware │
│ topics / services / TF │
├──────────────────────────────┤
│ Device communication │
│ UART / SPI / I²C / CAN │
├──────────────────────────────┤
│ Electrical signals │
│ scope / logic analyzer │
├──────────────────────────────┤
│ Physical hardware │
│ power / wiring / mechanics │
└──────────────────────────────┘
When one layer looks correct, move one layer down or up.
This prevents random debugging.
28. Debugging by Invariants
An extremely powerful technique is to define things that must be true.
Suppose a pipeline is:
camera
→ detector
→ tracker
→ planner
→ controller
Possible invariants:
camera FPS > 20
bounding boxes must lie inside image bounds
planner input timestamp must be recent
velocity <= hardware safety limit
If an invariant fails, you know which region of the pipeline is broken.
This is far better than simply saying:
Something looks weird.
29. Reproduction Is Half the Battle
A bug you can reproduce reliably is much easier to fix.
Try to reduce:
large system
into:
minimal failing case
For example, instead of debugging an entire API:
response = complex_pipeline(...)
reduce the issue to:
requests.post(url, json=minimal_payload)
If the failure remains, many unrelated components have been eliminated.
This is known as creating a minimal reproduction.
30. Binary Search Debugging
Git bisect is one example of a broader technique.
Suppose a pipeline has eight stages:
A → B → C → D → E → F → G → H
Output at H is wrong.
Check D.
If D is correct:
problem ∈ {E,F,G,H}
Check F.
If F is incorrect:
problem ∈ {E,F}
You have localized an eight-stage pipeline in only a few tests.
This debugging strategy appears everywhere.
31. Instrumentation
Sometimes existing tools cannot expose the information you need.
Then instrument the system yourself.
Add:
timers,
counters,
assertions,
request IDs,
memory statistics,
state snapshots,
health checks.
For example:
start = time.perf_counter()
result = model(input)
elapsed = time.perf_counter() - start
logger.info(
"inference_complete",
extra={"latency_ms": elapsed * 1000}
)
A well-instrumented system is dramatically easier to operate.
32. Assertions
Assertions document assumptions.
Example:
assert batch_size > 0
or:
assert tensor.shape[-1] == hidden_size
Instead of allowing invalid state to propagate through 50 functions, the program fails immediately near the source.
That improves debuggability.
33. Metrics, Logs, and Traces
These three are often called pillars of observability.
They answer different questions.
Metrics
How is the system behaving overall?
Examples:
requests/sec
GPU utilization
error rate
p99 latency
Logs
What happened?
Example:
request abc123 failed because upstream returned 503
Traces
Where did the request spend its time?
Example:
Gateway 5 ms
Database 8 ms
Inference 2.8 s
Together they provide a much more complete picture.
34. Performance Is Usually a Queueing Problem Somewhere
Imagine requests arriving faster than the server can process them.
Arrival rate:
λ\lambda
Service rate:
μ\mu
If:
λ≈μ\lambda \approx \mu
queues begin growing.
If:
λ>μ\lambda > \mu
the system cannot keep up indefinitely.
Then latency can explode even though individual computation time did not change much.
This is why production performance debugging requires understanding:
concurrency,
queues,
worker pools,
batching,
backpressure,
contention.
35. Concurrency Bugs
Concurrency creates especially difficult failures.
Examples:
race conditions,
deadlocks,
livelocks,
starvation.
35.1 Race condition
Two threads access shared state unpredictably.
Example:
Thread A reads x=5
Thread B reads x=5
Thread A writes 6
Thread B writes 6
Expected:
7
Actual:
6
Tools such as ThreadSanitizer help detect this.
35.2 Deadlock
Thread A holds lock 1 and waits for lock 2.
Thread B holds lock 2 and waits for lock 1.
A:
Lock1 → waiting Lock2
B:
Lock2 → waiting Lock1
Nobody can proceed.
Debugger thread dumps can reveal this state.
36. Performance Regression Debugging
Suppose release 1 processes:
1000 requests/s
Release 2 processes:
650 requests/s
Do not immediately optimize.
First establish:
regression=new version−baseline\text{regression} = \text{new version} - \text{baseline}
Then:
reproduce under controlled conditions,
compare profiles,
compare system metrics,
compare source changes,
bisect if necessary.
This is the performance equivalent of debugging correctness.
37. Hardware-Software Boundary Bugs
Some of the hardest bugs occur exactly where software meets hardware.
For example:
software says:
SPI write successful
but sensor returns nonsense.
Possible causes:
wrong SPI mode,
wrong clock polarity,
wrong chip-select timing,
voltage-level mismatch,
unstable power,
incorrect bit ordering.
The application sees only:
bad data
The logic analyzer sees:
MOSI bits incorrect
The oscilloscope sees:
clock edges malformed
Each tool exposes a different layer.
38. The Senior Engineer's Rule
Never ask only:
What do I think the problem is?
Also ask:
What observation would prove that hypothesis wrong?
Suppose you think:
The GPU is the bottleneck.
Test:
nvidia-smi
and GPU profiling.
If utilization is 10%, the evidence contradicts your hypothesis.
Good engineers prefer evidence over intuition.
39. A Practical Tool Selection Map
Think of the tools like diagnostic instruments.
Code changed?
↓
Git
Need to inspect remote system?
↓
SSH
Need process/system state?
↓
Linux tools
Need to find text/code quickly?
↓
grep / ripgrep
Build failing?
↓
CMake / build system diagnostics
Native crash?
↓
GDB
Memory corruption?
↓
sanitizers / Valgrind
Program slow?
↓
profiler
Request slow across services?
↓
tracing
Network suspicious?
↓
tcpdump / Wireshark
ROS behavior wrong?
↓
ROS graph / topics / TF
GPU slow?
↓
GPU profiler
Embedded communication wrong?
↓
logic analyzer
Electrical waveform suspicious?
↓
oscilloscope
Need objective comparison?
↓
benchmark
40. What “MASTER” Means Here
You do not need to memorize every flag of every tool.
Mastery means knowing:
What layer the tool observes
and:
What question it can answer.
For example:
If you hear:
segmentation fault
you should immediately think:
GDB
AddressSanitizer
core dump
If you hear:
request is intermittently slow
think:
metrics
tracing
logs
profiling
If you hear:
node exists but robot does not move
think:
ROS topics
TF
controller
driver
hardware
If you hear:
serial packets look corrupted
think:
logic analyzer
oscilloscope
If you hear:
it broke sometime last week
think:
git log
git diff
git bisect
That mental mapping is far more valuable than memorizing hundreds of commands.
41. The Complete Debugging Loop
A disciplined debugging process looks like:
Observe the failure\boxed{\text{Observe the failure}}
↓
Reproduce it\boxed{\text{Reproduce it}}
↓
Define expected behavior\boxed{\text{Define expected behavior}}
↓
Form hypotheses\boxed{\text{Form hypotheses}}
↓
Choose the right observation tool\boxed{\text{Choose the right observation tool}}
↓
Collect evidence\boxed{\text{Collect evidence}}
↓
Eliminate possibilities\boxed{\text{Eliminate possibilities}}
↓
Localize the cause\boxed{\text{Localize the cause}}
↓
Fix the cause\boxed{\text{Fix the cause}}
↓
Verify the fix\boxed{\text{Verify the fix}}
↓
Prevent recurrence\boxed{\text{Prevent recurrence}}
The last step may include:
automated tests,
assertions,
better logs,
monitoring,
benchmarks,
stronger types,
configuration validation.
A bug is not completely fixed if it can silently return tomorrow.
42. The Deeper Engineering Lesson
Debugging is not fundamentally about GDB.
It is not fundamentally about Git.
It is not fundamentally about Wireshark.
It is about observability.
Complex systems fail because their internal state differs from what we believe it to be.
Debugging tools let us compare:
our mental model\text{our mental model}
with:
reality\text{reality}
When those differ, reality wins.
The better your tools and mental models become, the faster you can discover exactly where that difference begins.
That is why senior engineers often seem unusually fast at debugging.
They do not necessarily type faster.
They do not necessarily know every command from memory.
They are better at asking:
Which layer could explain this symptom?
Then:
What is the cheapest experiment that would eliminate the most possibilities?
And then:
Which tool gives me direct evidence?
That mindset scales from:
a 30-line Python script
to:
a distributed AI platform
to:
an autonomous robot
to:
physical electronics.
The tools change.
The reasoning process does not.
Final Mental Model
When debugging, think vertically through the stack:
Source Code\boxed{\text{Source Code}}
Git, grep, compiler diagnostics.
↓
Program Execution\boxed{\text{Program Execution}}
debuggers, sanitizers, profilers.
↓
Operating System\boxed{\text{Operating System}}
processes, files, memory, sockets.
↓
Distributed System\boxed{\text{Distributed System}}
logs, metrics, traces.
↓
Network\boxed{\text{Network}}
packet capture, Wireshark.
↓
Accelerator\boxed{\text{Accelerator}}
GPU profiling.
↓
Robotics Middleware\boxed{\text{Robotics Middleware}}
ROS diagnostics, TF, topics, services.
↓
Device Communication\boxed{\text{Device Communication}}
UART, SPI, I²C, CAN.
↓
Electrical Reality\boxed{\text{Electrical Reality}}
oscilloscope, logic analyzer, multimeter.
↓
Physical Reality\boxed{\text{Physical Reality}}
power, cables, motors, sensors, mechanics.
A complete engineer eventually becomes comfortable moving through all of these layers.
Because when you build serious AI, robotics, and autonomous systems, the failure is not always:
“the Python code is wrong.”
Sometimes the model is waiting on the CPU.
Sometimes Linux killed the process.
Sometimes TCP retransmissions are destroying latency.
Sometimes the ROS TF tree is broken.
Sometimes CUDA is synchronizing unnecessarily.
Sometimes the motor driver is receiving no bytes.
And sometimes the wire simply is not carrying the voltage you thought it was.
The job of an engineer is to stop guessing and find out which one is actually true.