Taming a PoE switch with Home Assistant: a debugging story
The problem
I have a Netgear M4250-26G4F-PoE+ switch quietly powering a handful of devices over Ethernet. It’s a great switch with one habit I didn’t love: every PoE port defaults to ON at power-up. So each day my connected equipment would spring to life whether I wanted it to or not, and I’d log into the switch’s web UI to manually disable the ports I didn’t need. Every. Single. Day.
I already use Home Assistant to control AC smart plugs for other gear, with a tidy dashboard card: a master toggle at the top, individual toggles underneath, and the ability to flip everything on or off at once. I wanted the same experience for the PoE ports — except backed by SNMP instead of a smart plug.
The requirements
Written down, the goal looked like this:
- Auto-disable any PoE port that’s delivering power shortly after the switch powers on, so equipment doesn’t boot unattended. I’ll re-enable individual devices on demand.
- A Home Assistant dashboard mirroring my smart-plug card: master toggle, per-device toggles, live status, and a “Managed Devices” card showing only the gear that’s actually connected.
- Three visible states per port:
- delivering — device connected and powered,
- connected but off — device connected, but I’ve disabled the port,
- not connected — hidden entirely.
- Devices should be renameable right in the HA UI — no YAML editing to give a port a friendly name.
- Somewhat dynamic: newly connected devices should appear automatically, with manual add/remove as an acceptable fallback.
Simple enough on paper. The reality was a tour through several of Home Assistant’s sharper edges.
The design
SNMP under the hood
The M4250 exposes PoE control through the standard POWER-ETHERNET-MIB (RFC 3621), where N is the
port number:
- Admin state (
pethPsePortAdminEnable,1.3.6.1.2.1.105.1.1.1.3.1.N) — enable/disable a port. - Detection status (
pethPsePortDetectionStatus,1.3.6.1.2.1.105.1.1.1.6.1.N) — reportsdisabled,searching,deliveringPower,fault, and friends.
SNMPv3 with authentication and privacy (SHA-512 / AES-128) keeps the control plane locked down.
A two-layer switch per port
Home Assistant’s legacy SNMP switch platform can set the admin OID — but it has no
unique_id, which means you can’t rename it in the UI. The modern template switch
can be renamed, but doesn’t talk SNMP. So each port became a sandwich:
switch.poe_port_N_snmp— the legacy SNMP switch that does the real work.switch.poe_port_N— a template switch with aunique_idthat forwards on/off to the SNMP switch, mirrors its state, and is freely renameable.sensor.poe_port_N_status— an SNMP sensor on the detection OID, mapping the raw codes to friendly states.switch.all_poe_ports— a group over the template switches, giving me the master toggle.
The “managed list”
To show only connected devices (requirement #2 and #3), I needed to remember which
ports have real equipment behind them. That lives in a single input_text helper
holding a CSV of port numbers, e.g. 7,11,12,20. The dashboard’s auto-entities card
filters the full 24-port list down to just those.
Everything — switches, sensors, template entities, the group, the helpers, the automations, and the Lovelace dashboard — is emitted by one Python generator script, so the entire config is reproducible from a single source of truth.
The debugging story
This is where the blog post earns its title. Each fix revealed the next problem.
1. The template integration moved
On Home Assistant 2026.10, template switches must live under the top-level
template: key — the old switch: - platform: template form is rejected outright.
Easy fix, clear error message. A gentle warm-up.
2. The ghost _2 entities
After a config change, every entity reappeared as switch.poe_port_N_2, and the
dashboard showed “Entity not available” everywhere. The cause: Home Assistant’s
entity registry maps a unique_id to an entity_id permanently. A stale
registration still owned switch.poe_port_N, so the new entities were shoved to a
_2 suffix. The fix was to bump the unique_id so HA would mint fresh records that
claimed the clean names. (This was compounded by the HA disk being full, so registry
writes weren’t even persisting — a red herring that cost real time.)
3. The switch forgot its own credentials
Suddenly every sensor read unknown and the toggles went dead. The logs were full
of SNMP error: Unknown USM user. The culprit wasn’t Home Assistant at all — the
M4250 had lost its SNMPv3 user across a power cycle. On these switches, the SNMP
user (and the PoE admin state) live in the running config; if you don’t explicitly
save to startup config, a power-off wipes them. Lesson filed, save button
pressed.
4. The TupleWrapper that couldn’t .split()
With SNMP healthy, the power-on scan still refused to disable anything, logging:
UndefinedError: 'TupleWrapper' object has no attribute 'split'
This one is subtle and worth remembering. Home Assistant’s automation variables:
step applies native-type parsing to template results. My template built a CSV
string like "7,11,12" — and HA helpfully parsed it into a tuple. (A single value
"7" became an int; an empty string stayed a string.) So when a later step called
live_ports.split(','), it was calling .split on a tuple. Boom.
The fix was to stop round-tripping through CSV inside the automation: keep the value
as a native list the whole way through (a list literal always parses back to a
list), iterate it directly, and only join it into a CSV at the exact moment I write
the input_text. A nice detail: cv.string unwraps the resulting template object
via its .render_result, so the stored value is still a clean 7,11,12.
5. The real bug: a wiped list
Everything looked fixed — until I pressed the scan button a second time and my entire Managed Devices list vanished.
This was the deepest issue, and it wasn’t syntax at all — it was a data-model
mistake. The managed list had been defined as “ports delivering power right
now.” But the scan’s whole job is to disable those ports — after which they read
disabled over SNMP, which is indistinguishable from an empty port. So the
moment the scan ran a second time, “what’s delivering” was empty, and it dutifully
overwrote the list with nothing.
The realization: the managed list isn’t a live snapshot — it’s persistent memory of which ports have connected devices. SNMP can’t re-derive that once ports are off, so the list must be additive. The scan was rewritten to compute three lists — what’s delivering now, what’s already managed, and their union — write the union back, and disable only the live ports. It adds and disables; it never removes. Removal is a deliberate, manual act.
The outcome
With the union model in place, the three moments that matter all behave correctly:
| Transition | What happens |
|---|---|
| Switch powers on (via the equipment’s AC switch) | Connected ports boot powered, get added to the managed list, and are disabled after a boot delay. |
| Home Assistant restarts | The input_text helper restores its saved value, so the list stays populated. Nothing is re-disabled. |
| Scan button pressed | Powered ports are added and disabled; if everything’s already off, the list is left untouched — no wipe. |
Newly connected devices still appear automatically (a separate automation unions them in when they start delivering power), and I can rename any device straight from the Home Assistant UI. The daily login-and-disable chore is gone.
I also validated the whole config against the exact Home Assistant version I run
(Core 2026.10.0) using a local hass --script check_config harness, plus live
template-render tests to prove the type behavior — not just “it loads,” but “it does
the right thing for zero, one, and many live ports.”
Lessons learned
- Home Assistant’s
variables:and service data apply native-type parsing. A template that returns"7,11,12"becomes a tuple;"7"becomes an int. If you need a string, don’t assume you have one — keep structured data as native lists and only stringify at the boundary.cv.stringunwrapping.render_resultis the detail that makes CSV storage work cleanly. - The entity registry maps
unique_id→entity_idpermanently. Renaming collisions and_2suffixes come from stale registrations. Changing theunique_idforces a clean record; you can’t just rename your way out. - On modern HA, template entities belong under the
template:key, not under a platform of another domain. - Know your switch’s config persistence model. The M4250 keeps SNMP users and PoE state in running-config; without an explicit save to startup-config, a power cycle erases them. “It worked yesterday” can literally mean the hardware forgot.
- Model the state, not the snapshot. The biggest bug wasn’t a typo — it was defining “managed devices” as “currently powered,” when the system fundamentally can’t observe connectivity once a port is disabled. When a sensor can’t distinguish two situations, you need persistent memory, and your writes must be additive rather than destructive.
- Validate behavior, not just syntax. “Config loads successfully” and “the automation does the right thing for one, many, and zero ports” are different claims. Rendering the actual templates against a live core caught issues a config check never would.
- Step back when every fix spawns a new break. The turning point came from re-examining the requirements and the data model instead of patching the next error. The symptom was a wiped list; the real problem was a wrong definition.