Troubleshooting
Something not working the way this section says it should? Start here: search this page for words from what you’re seeing (a status, an error message, a button name), or jump straight to the area it’s under.
Before you dig in
For raw logs on the master, two services cover almost everything:
sudo journalctl -u buddy-server -f # the panel and its management serversudo journalctl -u buddy-agent -f # the co-located node agentWhen you need help beyond what’s here, export a diagnostics bundle from the panel (the Diagnostics button) — it captures versions, your node list, recent tasks and their logs, and cluster status, with no secrets in it.
Panel & access
- The master isn’t reachable at all after flashing (no SSH, no panel)
- The panel won’t open
- I don’t know the admin password
- I lost access to my authenticator app
- Every login says “Too many incorrect codes”
Setup and joining nodes
A node offline or NotReady
- A node shows as offline / NotReady
- The Activity panel no longer shows routine health checks
- Tasks keep failing, then work when you retry them
- Removal seems stuck
Network & overlay
Kubernetes health & recovery
Storage & apps
- A disk doesn’t appear / a “Prepare” option is missing
- A pool shows “degraded”
- “Detach” / “Delete” is refused
- Preparing a disk asks for confirmation, and there’s no undo
- “Prepare” fails with a filesystem error
- A package install fails with “cluster unreachable”
- Where does my app’s data live, and does it move with the node?
Migrating to an SSD
- “Migrate OS to SSD” is missing or blocked
- The node didn’t come back booted from the NVMe
- Boot order didn’t change / it still boots from the microSD
- “Retire microSD” is refused
- I want to undo a migration
Maintaining a node
- A node didn’t come back Ready after a reboot/update
- The action failed and the node is cordoned
- The drain is stuck
- A halted node won’t come back
- “Update firmware” is missing
- The panel went down while maintaining the master
- An Activity row for a reboot or shutdown looks wrong
Upgrading Kubernetes
- The upgrade is refused before it starts
- “An agent is ahead of the server”
- The panel blipped when the server upgraded
- A node is stuck at the old version / the roll halted
- The drain is stuck mid-roll
- The server upgrade itself failed (recovering with the recovery snapshot)
- “A cluster upgrade is already in progress”
Updating buddy-agent
- A worker is stuck on the old agent version
- “SHA-256 mismatch” in the logs
- A node needs an SSH re-bootstrap
- The roll is halted or “pinned” and won’t proceed
- “The master’s agent is not at the target version”
- A row is badged “same version, different build.”
- “An agent self-update is already in progress”
- An Activity row for an agent update looks stuck or frozen
Publishing on the Internet
- A Route shows a DNS sync error
- Port-forward shows “Not reachable yet” / readiness never confirms
- Behind CGNAT — use the tunnel
- A DDNS-only provider can’t get a certificate behind CGNAT
- “Remove the tunnel” shows “tunnel may still exist at Cloudflare”
- “Remove the Cloudflare token” fails outright
- A certificate stays “not issued” or shows an ACME error
- A Route shows “no pods running”
- Publish is stuck on “Continue” / the service picker