Homelab

Omada Link Backup: the standby WAN shows offline because nobody is checking it

TL;DR: On a TP-Link Omada gateway (ER707-M2, controller 6.2.14), a WAN port that is
the backup link in Link Backup mode reports Online Detection = offline (onlineDetection: 0)
even when the line behind it works fine. The gateway doesn’t seem to probe a standby
backup WAN at all. The moment the same port became a normal WAN again, it read online
with latency and loss figures within the same poll. So on a backup WAN, “offline” means
“not being checked”, not “down”. Don’t alert on it, and don’t trust it as proof that
your failover works.

The setup

Three internet lines on one Omada gateway:

WANLineRole
WAN1Fiber, ~930 MbpsPrimary
WAN5Second fiber, ~480 MbpsPrimary
WAN34G LTE through a MikroTik router, ~20 Mbps, meteredLast resort only

The LTE line should carry nothing unless both fibers are down. Omada does that with
Link Backup: primary WANs = WAN1 + WAN5, backup WAN = WAN3, and
“Enable backup link when all primary WANs fail”. Through the API it looks like this:

"wanLoadBalance": {
  "weight": "10,0,0",
  "linkBackup": true,
  "method": "backup",
  "mode": 1,
  "primaryWan": ["<WAN1>", "<WAN5>"],
  "backupWan": "<WAN3>"
}

One more detail matters here. The LTE router isn’t plugged straight into the gateway. Its traffic rides a
VLAN across two switches, and a patch cable hands it to the gateway’s WAN3 port. So WAN3
has link whenever the switch is up, even if the LTE modem is dead.
Link state tells you
nothing. I needed the gateway’s Online Detection (it pings or DNS-checks internet hosts
through each WAN) to be the thing that decides whether LTE is usable.

What I saw

The gateway detail in the controller’s API has a status block per WAN port. I read it on
four different days:

FieldWAN1WAN5WAN3 (backup)
status (link)111
internetState111
IP addressyesyesyes, by DHCP from the LTE router
onlineDetection110, every time for 10 days

The two primaries were always online. The backup was always offline.

My first reading was wrong

On the first day this looked like good news. The LTE modem had no cell signal that
afternoon (the router’s own log said failed to register on network), so a DNS lookup through it
returned SERVFAIL. WAN3 had link, it had an IP, and it still said offline. I wrote down:
“The gateway doesn’t trust link state. It won’t fail over onto a dead LTE line.”

The next day the modem found signal. Lookups of random names that can’t be cached
came back NOERROR / NXDOMAIN from the real upstream, so LTE was working. WAN3
still read onlineDetection: 0. It kept reading 0 on every check after that, while the line worked.

So the 0 on day one wasn’t the gateway spotting a dead line. It would have said 0 either
way.

That false “offline” spread further than one note:

  • My WAN dashboard graphs onlineDetection per line. LTE sat at “offline” on the graph the
    whole time
    while it was fine.
  • A watchdog design listed its LTE phase as “blocked: the LTE line has link but no
    internet.”
    That was a wrong conclusion drawn from a flag that wasn’t measuring anything.

The accidental experiment

Nobody set out to test it. On 2026-09-26 the WAN port settings were edited and saved in the Omada web UI.
A side effect I’ve written up separately: saving a WAN port in the UI re-sends the whole
load-balance object, and it came back without the Link Backup settings.
linkBackup went
true → false, and method, mode, primaryWan and backupWan disappeared.

In the same change, WAN3’s onlineDetection flipped from 0 to 1. Two days later it still reads:

2.5G WAN1  status 1  internetState 1  onlineDetection 1  latency 4 ms   loss 0%
WAN/LAN3   status 1  internetState 1  onlineDetection 1  latency 25 ms  loss 0%
WAN/LAN5   status 1  internetState 1  onlineDetection 1  latency 2 ms   loss 0%

Same port, same cable, same LTE line. The only thing that changed is that it stopped being
a standby backup and became a normal WAN (at weight 0). The gateway started
reporting latency and loss for it right away, so it’s clearly probing it now.

What that means

The most likely explanation: in Link Backup mode, the Omada gateway doesn’t run Online
Detection on the standby backup WAN.
It probes the ports it’s using, and the backup
just keeps its default “offline” until it’s needed. I haven’t found that written down
anywhere. It’s what the numbers show on this gateway and firmware.

It makes some sense. On a metered LTE line, pinging out every few seconds all month
costs data for nothing. That’s my guess at the reason, not something TP-Link says.

What it means in practice:

  1. onlineDetection: 0 on a backup WAN is ambiguous. It could mean “the line is dead”
    or “the gateway isn’t looking”. You can’t tell them apart from the controller.
  2. Don’t alert on it. Leave the backup WAN out of any “WAN offline” alert, or you get an
    alert that is always on and gets ignored.
  3. Check the backup line some other way. Ask the backup router itself (its modem
    status, or a DNS lookup of a random name through it, which can’t come from its cache), or
    probe out through it from a host that routes only that way.
  4. The open question is the one that matters: when both primaries die, does the gateway
    probe the backup before moving traffic onto it? If it does, a dead LTE line is handled
    correctly. If it switches on link state alone, it fails over onto a line that doesn’t work,
    and with my switch-in-the-middle wiring, link is always up. Only a real test answers this:
    unplug both primaries for a few minutes and watch.
    I haven’t done that yet, so the
    failover is still unproven.

How to check yours

With the controller’s internal API (the one the web UI uses; a Viewer account is enough):

GET /{omadacId}/api/v2/sites/{siteId}/gateways/{gatewayMac}
  → result.portStats[] → name, status, internetState, onlineDetection, latency, loss

GET /{omadacId}/api/v2/sites/{siteId}/setting/wan/networks
  → result.wanLoadBalance → linkBackup, method, mode, primaryWan, backupWan

If linkBackup is true and the port listed in backupWan shows onlineDetection: 0 with no
latency figure, you’re looking at the same thing I am. Check the line some other way before
you decide it’s down.

And after any edit to WAN settings in the web UI, GET wanLoadBalance again and
compare. The Link Backup settings can silently vanish on save, and all that happens is that the backup
port starts looking healthier.

What I learned beyond Omada

  • A health flag you haven’t seen change doesn’t tell you much. Ten days of “offline” felt
    like data. It was a default. Before trusting a status, see it flip both ways for a real reason.
  • A result that agrees with you still needs a second test. Day one’s dead modem made the
    flag look right, and I wrote down a conclusion that turned out to be a coincidence.
  • Accidents are experiments too. The unwanted UI save was the only thing that changed
    one variable, and it gave the clearest evidence I have. Write down everything that changed in
    the same moment, not only what broke.