BGP PIC vs ECMP: What Your Backup Paths Cost the Forwarding Chip
Everyone enables backup paths and ECMP for resilience. Almost nobody checks what it costs the forwarding chip until the logs start screaming
We recently watched a Jericho-based edge run out of egress encap space. Label-imposition rewrites plus BGP backup paths pinned a hardware bank at 100%, and routes started failing to program.
It all traces back to a common design choice. You receive more than one path for a prefix. What should the hardware do with them?
Option 1: Primary + backup (BGP PIC Edge)
One path forwards. The second sits pre-programmed in the FIB. When the primary next-hop dies, the linecard swaps locally with no BGP reconvergence. Failover lands in tens of milliseconds, whatever the table size.
Option 2: ECMP across all paths
Every equal-cost path forwards at once. Traffic load-shares, and if one path drops, flows rehash onto the survivors.
What matters in production:
PIC main + backup
+ Fast, deterministic failover
+ Works when paths are NOT equal-cost. This is the big one.
+ Lighter state: one active path plus one backup
– Backup capacity sits idle until something breaks
– You trust best-path selection to pick the right primary
ECMP
+ Uses all your links
+ No single primary to worry about
– Needs truly equal paths: same local-pref, AS-path length, MED, and IGP metric to the next-hop. Miss one and you silently fall back to a single path
– One member failure can reshuffle flows across every member unless you run resilient hashing
– More forwarding state per prefix
How do you run it on your edge: all-paths ECMP, PIC backup, or a mix?
The part people forget is that both cost hardware. Every backup path and every extra ECMP member burns egress resources on the ASIC: FEC entries, next-hop groups, encap rewrites. At full-table scale, “backup for every prefix” or “ECMP everything” adds up fast.
Rule of thumb:
Exits not equal-cost and you need fast failover? PIC backup.
Paths equal-cost and you want the capacity? ECMP.
Doing either at full-table scale? Check NPU resource usage before you enable it network-wide, not after the logs start screaming.
Most real designs end up using both. ECMP where paths are equal, PIC backup protecting whatever set you install.







