SNMP discovery¶
Monitoring tells you whether a device is up. SNMP discovery tells you what the device actually is right now - its system facts, its interfaces, its neighbours - read straight off the box over SNMP into a read-only observed layer. Danbyte stays the source of truth; discovery never silently overwrites your intended configuration. When observed reality and intent disagree, that difference surfaces as drift you can review and accept one item at a time.
This page is organised by task. Jump to:
- The observed-vs-intended model - why discovery is safe
- SNMP profiles - reusable v1/v2c/v3 credentials
- Credential hierarchy - device → role → type → location → site → default
- Poll a device - read system facts + interfaces
- Scheduled polling & utilisation - the sparkline series
- Drift & reconciliation - accept observed into intent
- Topology: LLDP & ARP - neighbours and the ARP table
- MAC tables - learned MACs per port, uplinks, history and Refresh MACs
- Custom SNMP sensors - vendor health OIDs, and sharing them as a pack
- Permissions
Observed vs intended¶
Everything SNMP reads lands in a separate observed store
(DeviceSnmp), never on the Device source-of-truth fields. So a poll can run a
hundred times and your device record is untouched. The only place observed data
flows back into intent is when you explicitly accept a drift item - a
deliberate, permission-gated click. That's the whole design: reality flows in on
demand, but you decide what becomes truth.
SNMP profiles¶
A profile is a reusable set of SNMP credentials, named per tenant. Manage them under Settings → SNMP profiles.
- Version -
v1,v2c, orv3. - v2c - a community string.
- v3 - username + auth/priv protocols and keys.
Secrets (the community, the v3 keys) are encrypted at rest and write-only
over the API - a GET never returns them, only a has_secrets flag. This
mirrors how monitoring check credentials are stored. An edit changes only the
keys it sends - a new auth key keeps the stored privacy key - and a key sent
as null is removed.
Mark one profile default for the tenant. Setting a new default automatically clears the previous one, so there's always at most one default and switching it actually switches it.
Per-VLAN MAC tables (Auto / Always / Off) says whether the agent's
per-VLAN forwarding tables are read - see MAC tables. Saving
the form keeps the parameters it has no field for (mac_max_vlans,
mac_budget_s, a port set over the API).
Credential hierarchy¶
You rarely want to pick a profile per device. Instead, bind a profile at the level that makes sense and let it inherit. When Danbyte polls a device it resolves the effective profile most-specific-first:
- Device - a profile bound directly to this device.
- Device role - e.g. all
core-switchdevices. - Device type - e.g. all
C9300-48P. - Location - bound on the device's location, inherited down from a parent location if the child doesn't set one.
- Site - bound on the device's site.
- Tenant default - the profile flagged default.
Levels 1–3 are what a device is; levels 4–5 are where it lives. The location/site levels let a remote Outpost poll a site's devices with site-scoped credentials - set them on the site or location form. If nothing is bound and there's no default, Danbyte will only auto-pick when the tenant has exactly one profile - otherwise it declines rather than guess which credential to poll with. The device's SNMP card shows where the effective credential came from.
Virtual routers & appliances¶
SNMP polling isn't only for physical devices - a virtual machine running a router/firewall OS (MikroTik CHR, VyOS, pfSense, …) can be polled too. Open a VM → its SNMP tab → Poll now. Danbyte reads the same system group + interface tables over the VM's primary IP and stores them on the same observed store, so facts, interfaces, LLDP neighbours and ARP show exactly as they do for a device.
Profile resolution follows the equivalent hierarchy for what a VM is and where
it runs: VM → platform → cluster → site → tenant default. Bind a profile
at whichever level fits (a per-VM binding can also carry a target override to
poll a management address instead of the primary IP).
Poll a device¶
Open a device → its SNMP tab → the Observed card → Poll now.
Danbyte does one synchronous SNMP read of the system group (sysName,
sysDescr, sysObjectID, sysUpTime, sysContact, sysLocation) plus the
interface tables (ifTable/ifXTable), and stores them as observed facts. The
card shows a reachable / unreachable badge, the named facts (never raw
OIDs), the interface list with oper-status and speed, and the last-polled
timestamp.
A poll never touches the device's source-of-truth fields - it only refreshes this card.
Poll now also reads the device's MAC table, but quickly: it has to answer before the web server gives up on the request, so it stops after a short time budget and skips per-VLAN tables. A big switch can come back with a partial table that way; Refresh MACs reads the whole thing in the background.
The tab is laid out by content width: the system facts, the interface table and the drift inbox run full width, and the narrow cards below them - LLDP neighbours and the ARP table, custom SNMP sensors and the BMC - pair up two-across on a wide window and stack on a narrow one.
Polling a stack¶
A virtual chassis answers SNMP as one box: whichever member you poll, the agent reports every member's ports plus the stack's logical interfaces (port-channels, VLAN and loopback interfaces, the management port). Danbyte therefore polls a stack once, through its owner - the designated master, else the lowest-positioned member - and stores the observation on that member. Polling any member, on the device page or on the schedule, polls the owner; a member without an address of its own is reached through the owner's. The member's Observed card says Polled via stack member ….
The observation is then split back onto the members it describes, so each member's drift and Sync from SNMP only ever see its own slice:
- a name that matches an interface the member already has, or that its
device-type or module templates render for its position
(
{position}), belongs to that member; - otherwise the first number after the leading letters names the member
slot -
Gi2/0/1,Ten-GigabitEthernet2/0/1,ge-1/0/0,1/1/1; - everything else -
Port-channel1,Bridge-Aggregation1,Vlan1,Loopback0, the management port - belongs to the owner.
A port is never proposed as new on one member while another member already has it, and a logical interface that lives on the master is never stale on a member. Where a vendor's naming defeats rule 2, the port lands on the owner: open it, and the Stack member field on the interface form moves it - the cable, IPs and MAC objects follow.
The stack page has its own SNMP tab with Poll stack, Sync stack from SNMP (each member in position order) and every member's drift inbox; the fleet drift view lists one row per member.
The stack's MAC table is recorded once, from the owner's poll, with each learned MAC placed on the member that owns its port by the same rules. An Outpost that polls every member adds nothing a second time.
Scheduled polling & utilisation¶
The on-demand button is a snapshot. To build a utilisation series for the per-interface sparklines, run the poller on a schedule:
Each run is a full poll, MAC table included, and records the interface HC
octet counters (ifHCInOctets / ifHCOutOctets) as a time-stamped sample. Utilisation is then derived as a rate
between consecutive samples - Δoctets · 8 / Δt, as a percentage of the
interface speed. A counter that goes backwards (reset/reboot/wrap) yields a 0
delta rather than a negative spike. Schedule poll_snmp from cron or a systemd
timer at whatever interval you want the sparklines sampled.
Samples are kept for MONITORING_SNMP_SAMPLE_RETENTION_DAYS (3 by default);
the daily monitoring prune deletes older ones, and the sparklines read only
that window. Before 0.17 nothing pruned them, so an install that polled for a
long time sheds its backlog on the first prune after the upgrade.
Counter64-safe
HC octet counters are SNMP Counter64 (unsigned 64-bit). Danbyte stores them as a 20-digit decimal so a large counter on a long-running, high-traffic interface can't overflow and crash the poll.
Hardware health runs itself. The danbyte-hardware systemd timer polls
every configured BMC (Redfish) and custom SNMP sensor
every 30 minutes, reconciling inventory and flipping statuses - so a
failing disk turns red on its own, no button press. Scheduled scope is bounded
to devices with a Redfish endpoint or a device-type-scoped sensor (plus,
when a tenant has an all-types sensor, every device with a primary IP);
all-types sensors otherwise run on the device's on-demand Poll sensors
button. Run it by hand with python manage.py poll_hardware.
Drift & reconciliation¶
The drift inbox on the device page compares observed state to your intended configuration and lists the differences. Usually that observation is Danbyte's own SNMP poll; an integration that already watches the device can offer one too (see Zabbix), and where both speak to the same field the poll wins - walking the device is better evidence than a second-hand account of it. An item raised by anything other than the poll is labelled with the source that raised it.
The differences:
- Device name vs
sysName. - Interface present on the device but not in Danbyte (
interface_missing). Accepting it - or Sync from SNMP - creates a port the agent reports as a loopback, propVirtual, tunnel or VLAN interface (l3vlan,l2vlan) with type Virtual, so it is virtual: off the faceplate and out of port utilization by default. Aggregates arrive typed LAG (below); everything else arrives with no type. - MAC, admin-status, VLAN or speed mismatch on an interface you already have.
- Stale - Danbyte has an interface the device no longer reports (shown for awareness; discovery never deletes from the SoT).
- LAG membership (
lag_membership) - the aggregate a port reports itself under differs from its LAG / aggregate in Danbyte. See Link aggregation.
Changed in 0.17
Discovery used to create loopbacks, SVIs, tunnels and VLAN interfaces with no type, as ordinary ports, so they counted in port utilization. The upgrade marks the existing ones virtual by the type their last poll reported; a type somebody set by hand is left alone.
Wherever a component is drawn, a difference shows as an amber outline next to the record rather than replacing it: on the photo faceplate, on the 3D room's port and bay markers, and as a drift pill in the component tables. Clicking the marker names the difference ("SNMP says failed", "speed: SNMP says 1 Gbps"). The device header carries a count badge, and the Components tab and its sub-tabs are dotted when something inside them differs - so you can find drift without opening every tab.
Where drift is marked¶
Drift is flagged in place, on the record it disagrees with, so you never have to open a device to learn that it differs:
- The device's Components → Interfaces table shows an amber drift badge on each affected row, and clicking it lists exactly what differs on that port.
- The fleet Interfaces list marks every drifted port with a quiet amber compare-arrows glyph next to its name. The tooltip counts the differences and names their kinds (interface mismatch, not reported by SNMP, IP not recorded on this port), and the marker links to that port's device → Components → Interfaces. The whole-stack interfaces table on a virtual chassis marks its ports the same way, for every member at once.
- The Devices list carries the same glyph per device, next to the compliance violation triangle - a rule you wrote failing and the device reporting something else are different problems, so they get different marks.
Every marker is a read-only signal: reviewing and accepting drift stays in the drift inbox, so the source of truth only changes when you choose (Danbyte stays drift-aware, never drift-driven).
Prefixes and IP addresses have no drift marker
Deliberately - no drift item references a prefix or an IP address that already exists in Danbyte. A Discovered IP is an address SNMP saw that Danbyte doesn't record, so there is no row to mark; the item is reported against the interface it was observed on (and marked there), plus, when no prefix contains it, as a prefix to add. Prefixes and recorded IPs are never compared to observed state, so a marker on those lists would always be blank.
Excluding a port from drift¶
Some ports can never be polled - the silkscreened host NICs a BMC agent doesn't see, an out-of-band jack, a port on gear behind the managed device. Left alone they flag as Stale - not seen on device after every poll, forever, and dismissing only hides them until the next one.
Click Exclude on the stale row instead. It sets the interface's Exclude from SNMP drift flag: the port stops being compared in both directions - never reported stale, never mismatch-checked, never touched by Sync from SNMP - while everything else about it (cables, IPs, monitoring) behaves as normal. Excluded ports show a muted eye-off mark in the interfaces table, and the flag is a checkbox on the interface's edit form, which is also where you undo it.
MAC comparison is separator-insensitive - 00:11:22:33:44:55 and the Cisco
dotted form 0011.2233.4455 are recognised as the same address, so reformatting
alone never shows as drift.
Click Accept on an item to write that observed value into intent. This is the
only action that mutates the source of truth, and it requires the same
device.change permission the device form does (see
Permissions). Everything else on this feature is read-only.
Drift kinds:
- Device name -
sysNamevs the device name. - Serial - what an integration's inventory reports vs the device's serial. Danbyte's own SNMP poll does not read a serial, so this one only ever comes from a source that does.
- New interface - observed on the device, missing in Danbyte.
- Interface mismatch - MAC, admin-status, VLAN or speed differs. Speed
is compared as a number, so
1G,1 Gbpsand an observed 1000 Mbps are the same value - reformatting never reads as drift, and an intended speed that isn't parseable ("dual 10/25") is treated as deliberate and left alone. - Discovered IP - an IP SNMP sees on an interface that Danbyte doesn't record. Accepting it assigns the IP to that interface (binding an existing unassigned IP if one matches, otherwise creating it in the smallest containing prefix). It then appears on the device's IPs tab - closing the discover→assign loop. If no prefix contains the address, accept fails: add the prefix first.
Linking a discovered name to a port you already made¶
The names on the silkscreen and the names the agent reports rarely match: you
labelled the port Ethernet 1, the switch reports it as eth0. Discovery sees
two things where there is one, and the pair drifts forever as both new and
missing.
On any New interface row, Link to… lists the device's own interfaces (unlinked ones first, searchable) - pick the port that discovered name really is. Danbyte stores it as the interface's SNMP name, the matcher starts treating the two as one, and both drift rows disappear on the next poll.
Linked ports carry an ↔ eth0 badge next to their name in the interfaces
table, so a link is never invisible.
To remove a link, click the ↔ badge on the port and choose Unlink -
the undo sits on the thing it undoes. The interface form's SNMP name field
does the same job if you're already editing the port. Linking a name that
another port has already linked moves it; a discovered name belongs to exactly
one port.
A link replaces the port's label rather than adding an alias to it. Saying
"the agent calls this port eth0" also says the agent never reports
Ethernet 1, so Danbyte stops expecting the label - otherwise the port you just
linked would keep drifting as not seen on device forever.
You can't link onto a name another port already has
If eth0 exists as an interface in its own right, IMM cannot be linked to
eth0: both would answer to that name, and only one can win the match. The
duplicate is the actual problem - delete or rename the port you don't want,
then link. Danbyte refuses the link and says so rather than accepting one
that can't work.
Link aggregation¶
The agent reads the bundle a port belongs to from IEEE8023-LAG-MIB
(dot3adAggPortAttachedAggID), falling back to IF-MIB's ifStackTable where a
port stacks under an aggregate interface. Every observed interface row then
carries lag_if_index - the aggregate's ifIndex, blank when the port is not a
member - and an aggregate reports type_name: lag even where the box calls it
propVirtual.
What that does in the inbox:
- A new interface row for an aggregate carries a
LAGbadge; accepting it creates the interface with type LAG (so it can take members). - A LAG member row shows
Gi0/1 Po1 → Po2(or- → Po1for a port that joined a bundle,Po1 → -for one that left). Accept sets - or clears - the port's LAG / aggregate. If the aggregate does not exist here yet the row says accept Po1 first and cannot be applied until it does. - Membership is compared by aggregate name, so a stack reports the master's
Po1on every member without false drift, and accept resolves the aggregate across the virtual chassis. - Update only still reports membership - it is a field on a port you already have, not a new port.
- An aggregate created before types were enforced (blank or Virtual) is promoted to type LAG on accept; one typed as physical media is refused - fix its type first.
- Sync from SNMP creates missing aggregates typed LAG and applies
memberships after the interface pass (
lag_membershipsin the summary).
An Outpost older than this feature never sends lag_if_index; its devices show
no membership drift until the agent is updated (see
Outposts).
Sync from SNMP¶
The drift inbox accepts items one at a time. The Sync from SNMP button on the
device's Interfaces tab does it all at once: create every observed interface
Danbyte lacks, fix MAC / admin-status / speed / VLAN drift, and assign
every observed IP that has a containing prefix. It reports what it
created/assigned and how many IPs were skipped for want of a prefix. (The device
name is left alone - accept that explicitly.) Needs device.change.
What a poll/sync reads per interface:
- Speed -
ifHighSpeed→ "10 Gbps" / "100 Mbps". - Layer - L3 if the interface has an IP (
ipAddrTable), else L2. - Access VLAN - the PVID from Q-BRIDGE-MIB (
dot1qPvid, mapped to the ifIndex via the bridge-port table), with the name fromdot1qVlanStaticName. On sync the VLAN becomes a first-class Danbyte VLAN object (find-or-create, ungrouped) and is assigned to the interface. L3-only devices and non-switches don't report it - that's fine.
Loopback and other special addresses
Observed addresses that don't belong in IPAM - loopback (127.x, ::1),
link-local (169.254.x, fe80::), unspecified (0.0.0.0, ::) and
multicast - are recognised by range and never offered for import or flagged
as drift, even though the Observed card still shows them as the device
reports them.
Fleet-wide drift view¶
The per-device card is for one box. To see drift across the whole fleet, open Drift in the sidebar - it has two tabs:
- Config (Ansible) - config-drift reported by your runner (device config vs rendered template).
- SNMP (observed) - every SNMP-polled device with its drift status (in sync / N drifted / unreachable), a one-line summary of what drifted (name, interfaces), the profile used, and when it was last polled. Filter by status; click a device to open its drift inbox and accept items.
Both tabs answer the same question - does reality match intent? - from the two sources Danbyte has (your runner, and SNMP).
Topology: LLDP & ARP¶
A poll also walks LLDP-MIB for directly-connected neighbours and reads the device's ARP table. The device's SNMP tab renders both as their own cards, side by side below the interface table:
- LLDP neighbours -
local-port ↔ remote-device : remote-port. - ARP table - the IP ↔ MAC pairs the device has learned.
Both are three narrow columns, so they pair up rather than stretch across the page; a device that reports neither simply doesn't show them.
The join logic (parse_lldp / parse_arp) is pure and unit-tested, so it's
correct independent of any one device's quirks.
Switch-link suggestions & the uplink guard¶
On a bridging device, drift suggests which access port each already-tracked IP hangs off - reviewed and accepted like any other drift, and applied by Sync from SNMP. Since 0.17 the suggestion follows the MAC's Location: an IP is suggested on a port of this device when its MAC is located there as an access sighting. Location is one answer for the whole network, so two switches can no longer both claim a host and re-claim it from each other on every poll.
The IP comes from the ARP table of any polled device - a router, an L3 switch, a firewall, a virtual router - with the switch's own table asked first, then the others in device-name order; the first answer per MAC wins. ARP sources (Settings → Monitoring → Switch-link suggestions) narrows that to the devices you name, in name order, exactly as before: on L2-only networks list the gateways and firewalls that actually route. If two sources disagree about a MAC, the first answer in device-name order wins, deterministically, rather than flapping between polls.
Uplinks never get suggestions: a trunk learns every MAC behind it. The old fixed limit of four MACs is now the tenant's Uplink above setting, and Uplink: Never on an interface lets a busy port take suggestions anyway. A device not polled since the upgrade keeps answering from its pre-0.17 tables, with the same rules, until its next poll.
Ghost cables on the topology map¶
LLDP also feeds the topology map (/topology). Real cables render as solid
edges; where two devices are LLDP-adjacent but have no cable in Danbyte, a
dashed ghost edge appears (and a "N LLDP links" chip in the header). LLDP
neighbours are matched to devices by name or observed sysName, so links show
up even before you've reconciled a name.
A virtual chassis answers as one box: every
member's table lists the whole stack's neighbours, and a neighbour names the
stack rather than a member. Each link lands once, on the member that owns the
port - the one with an interface of that name, else the one whose slot the
name carries (Gi2/0/1 is member 2), else the stack's master - so a polled
stack draws one ghost per link, not one per member.
Click a ghost edge to materialise it into a real Cable. SNMP can't report
the physical connector, so you pick the cable type (and, if the devices are
adjacent on more than one link, which port pair). Creating the cable needs
cable.add, and both interfaces must already exist - if an end is missing,
accept its interface drift first. Once cabled, the ghost is replaced by a solid
edge.
MAC tables¶
A switch's forwarding table says which MAC addresses it learned on which port. Every poll of a bridging device reads it, and Danbyte keeps it as sightings: one MAC, on one port, in one VLAN, with when it was first and last seen and when it went away.
What is read and kept¶
The table comes from the 802.1Q (Q-BRIDGE-MIB) forwarding database, falling
back to the plain BRIDGE-MIB one. Some agents (Cisco among them) keep one
table per VLAN; the SNMP profile's Per-VLAN MAC tables option decides
whether those are read: Auto (default) reads them when the agent lists
them, Always reads them anyway - using the VLANs Danbyte has on the
device's ports when the agent names none - and Off never does. Over the
API the option is the profile's params.mac_vlan_contexts
(auto/always/off), with mac_max_vlans (1-1024, default 128) and
mac_budget_s (5-600 seconds, default 120) next to it.
Only learned entries are kept. Dropped: the switch's own and invalid entries, multicast and broadcast addresses, all-zero MACs, and any MAC that is one of the polled device's own interface addresses - so a port never "learns" itself. Entries marked mgmt or other (port security, static entries) stay.
Each entry is placed on the stack member and the Danbyte interface that own its port, by the same name matching drift uses - the agent's ifName or ifDescr, or an SNMP name link. A port Danbyte doesn't have yet keeps its observed name; the poll after you add it links it. A MAC learned on a port-channel stays on the aggregate. The VLAN is the one the agent reports; it is blank where the agent shares one table across VLANs, and for an agent that predates MAC tracking.
ARP tables are kept the same way: one IP ↔ MAC pair per device (or VM), with first and last seen, which is what joins a MAC to its IP address. An ARP read closes entries only when the agent says it finished; one from an agent that predates MAC tracking counts as finished when it returned any rows.
Complete and partial reads¶
A read that finished is the truth: MACs it no longer reports are closed
as gone. A read that stopped early - the time budget ran out, the
100,000-row cap was hit, a walk failed, per-VLAN tables could not be opened -
is partial: it adds and refreshes what it saw and closes nothing, so a
slow switch never reads as "every MAC left". VLANs left out of a read - over
mac_max_vlans, or a per-VLAN table that failed - don't make it partial, so a
big switch still closes what moved elsewhere; the MACs in those VLANs are
simply neither closed nor refreshed by that read. A poll that never reached
the device writes nothing at all.
The device's SNMP state carries fdb_polled_at - the last complete read - and
fdb_meta, how the last read went: the source table, complete,
truncated, how VLANs and ports were mapped, what was dropped and why, and
the error, if any. Credentials never appear in it.
History and retention¶
A sighting is open while the MAC is there. A MAC that moves from Gi1/0/5
to Gi1/0/9 closes one row and opens another, and that pair is its history.
Rows unseen for longer than Forget MACs unseen for (Settings → Monitoring,
default 30 days) are closed and then deleted by the daily prune - so a switch
that stops answering stops locating MACs once its rows age out - and closed
rows are deleted once they are that old. A row last seen more than a day ago
reads as stale. ARP entries follow the same rules.
Privacy
A month of which MAC sat on which port with which IP address is close to personal data on an office network. The retention setting bounds how long it is kept, and every read follows the viewer's permissions.
Sightings are observed data that change on every poll: like DNS records they are not in the change log, send no webhooks and are not in the search index. No MAC objects are created for learned MACs; Add object on a MAC page still makes one by hand.
Uplinks¶
A trunk learns every MAC behind it, so an uplink must never be where a MAC "is". A port is an uplink when any of these holds:
- Its interface says Uplink: Always.
- Its LLDP neighbour is a switch: one whose announced capabilities include bridge or router but not telephone - an IP phone announces bridge + telephone and stays an access port - or a device Danbyte polls with a MAC table. LLDP switch neighbours mark uplinks turns this rule off.
- It is a LAG aggregate or a LAG member.
- It learns more distinct MACs, counted across VLANs, than Uplink above (default 4; 0 turns the count rule off).
Uplink: Never on the interface beats rules 2-4, and Always beats everything (see the interface form). The rules are evaluated when the table is read, from what the last poll stored and the current settings, so changing a setting applies without a re-poll. Every answer carries its reasons - LLDP neighbour sw-core-01, 6 MACs, above 4, Set on the interface, Aggregate.
An uplink's MACs are still stored: the port shows a count and lists them on demand, each with where it really sits. An uplink is never a MAC's Location while any switch reports the MAC on an access port, and it gets no switch-link suggestions.
Location¶
A MAC's Location is the port it really sits on: of its present sightings on devices you may view, an access port wins over an uplink; with several, a sighting that began after another was last seen replaces it (the MAC moved), one more than a day older than the newest is dropped (a switch that stopped answering), then the port with the fewest MACs, then the device and port name. When no switch reports it on an access port - a desk switch Danbyte doesn't poll, say - the Location falls back to the uplink with the fewest MACs, marked behind uplink, so the MAC is still found. See Where is this MAC?.
On the device's pages¶
On a device that reads a MAC table, the SNMP tab's interface table shows Learned MACs instead of the ports' own hardware addresses:
- Each port lists up to MACs shown per port MACs
(Settings → Monitoring), one per line with a
muted name · IP; a phone seen in its voice and its data VLAN is one line.
+N more opens the port's whole list - MAC, VLAN, IP, name, first seen. A
MAC the port learned but that sits elsewhere shows where
(
→ sw-acc-07 · Gi1/0/12). - An uplink shows an Uplink badge, its reasons in the tooltip, and a count instead of a list. Clicking the count lists the MACs seen through it and where each really sits, and links to the port's MACs tab.
- Under the table,
MAC table · 412 MACs on 37 ports · read 3m ago- the last complete read. A partial badge marks a read that stopped early; its tooltip says why. - Refresh MACs, beside Poll now, starts Refresh MACs and
reads
Refreshing…until the run is over; a toast then gives the count.
Components → Interfaces has the same Learned MACs column, an uplink
chip after the name of each uplink and Refresh MACs in its toolbar; the
whole-stack table reads the stack's MACs the same way. Only people who may
change the device see Refresh MACs.
Refresh MACs¶
Refresh MACs reads one device's whole MAC table - the full time budget and the per-VLAN tables - in the background, where Poll now has to be quick. On a stack member it refreshes the stack's owner. It needs change on the device, which the job checks again when it runs, so a permission revoked in the meantime stops it. One refresh per device runs at a time; asking again while one runs returns that run. A device an Outpost polls is queued for its Outpost, like Poll now. If the job queue is unavailable the refresh runs at once in the quick mode. There is no per-port refresh: SNMP cannot read one port's MACs without walking the whole table.
The API¶
| Endpoint | What it returns |
|---|---|
GET /api/monitoring/devices/<id>/macs/ |
Learned MACs per port: each port's uplink state and reasons, its MAC count, how many are located there, and up to "MACs shown per port" MACs with vendor, VLANs, IPs, name, first and last seen, and where each really sits. ?view=observed gives a stack owner's whole observation - the ports of the members you may view; ?limit= overrides the per-port count (0 = all). Device view. |
GET /api/monitoring/interfaces/<id>/macs/ |
One port's MACs, ?state=present (default) or all with the gone history, paged by ?cursor= and ?limit=, plus the port's uplink state. Interface view. |
GET /api/monitoring/mac-sightings/ |
The network-wide learned table, one row per MAC at its Location - see the Learned list. |
POST /api/monitoring/devices/<id>/mac-refresh/ |
Starts Refresh MACs: 202 {queued, run_id, running}, or {queued_on_outpost} for an Outpost's device. |
GET /api/monitoring/mac-refresh/<run_id>/ |
The run: status (queued, running, done, unreachable, denied, skipped, error), done, macs, ports, complete, error. For the user who started it, or anyone who may view the device. |
GET /api/monitoring/devices/<id>/snmp/ |
Now also fdb_polled_at and fdb_meta; the raw table is not returned. |
BMC hardware health (Redfish)¶
Servers expose their hardware over their BMC's Redfish API - the DMTF management standard that iDRAC (Dell), iLO (HPE), XClarity (Lenovo), Supermicro and Cisco UCS controllers all speak. Danbyte can poll it and keep the device's inventory items in sync - disks, CPUs, DIMMs, PSUs and fans, with real serials and live health.
Set it up on the device's SNMP tab → BMC (Redfish) card: enter the
BMC address, port and credentials (encrypted at rest, never returned by the
API), then Poll now. The collector walks
Systems → Storage/Processors/Memory and Chassis → Power/Thermal, and
reconciles what it finds:
- Parts are matched by serial number first, then by name - so renaming a
disk (e.g. to match a drawn
Bay 3marker) sticks across polls. - Missing parts are created with kind, media (NVMe/SSD/HDD), capacity and model; existing parts get their hardware facts updated. Nesting, tags, descriptions and custom fields are never touched.
- Health → status:
OK→ Active,Critical/Warning→ Failed - so a failing disk turns red on the Hardware tab, the photo faceplate and the 3D rack. Status flips are journaled on the device. Parts the BMC stops reporting are left alone.
BMCs live on management (RFC1918) networks, which Danbyte's outbound-request guard normally blocks. A Redfish endpoint is a deliberate, scoped exception: it's configured by someone with device-change permission, pinned to that one host, fetched with redirects disabled, and loopback/link-local addresses are still refused. TLS verification is off by default (BMC certificates are usually self-signed) - enable it when yours chain to a trusted CA.
Custom SNMP sensors (vendor health OIDs)¶
Not every BMC speaks Redfish - plenty are SNMP-only (Supermicro, older iDRAC/iLO, Synology, storage shelves). SNMP has no standard hardware- health MIB, so each vendor exposes disk/PSU/fan status under its own OIDs. Custom sensors let you teach Danbyte those OIDs.
Find the OID by looking, not by reading a MIB¶
You normally need the vendor's MIB file to know which OID reports health. Explore OIDs on the Custom SNMP sensors card removes that step: it walks down the device's own OID tree with you, one level at a time, and shows a table as the table it came from the moment you reach one.
Start anywhere - 1.3.6.1.4.1 (the root of every vendor's private tree) is
offered in the field. Each level lists its branches with the first value found
underneath, as a hint at what's down there:
Open one to go deeper. Danbyte recognises a table when every branch holds its values exactly one level down, and switches to a grid automatically - no need to know in advance whether you're looking at a branch or a table.
Why browsing isn't just a walk
A walk returns OIDs in lexicographic order, so walking 1.3.6.1.4.1
directly spends its whole budget inside the first vendor it meets and
never reveals the others. Browsing costs one request per branch instead of
one per value, which is what makes the vendor tree reachable at all.
Reaching a Lenovo IMM's power-supply table at
1.3.6.1.4.1.2.3.51.3.1.11.2.1:
| Row | .1 | .2 | .5 | .6 |
|---|---|---|---|---|
| 0 | 0 | Power System | Unknown | Normal |
| 1 | 1 | Power Supply 1 | K135155D0K2 | Normal |
| 2 | 2 | Power Supply 2 | K135155D0K5 | Normal |
Column .2 names the supplies, .5 holds serials, and .6 is health. Click
.6 → Create sensor, and the form opens with that column's OID filled in
and every value it returned already listed, so writing the value map is a
dropdown per value instead of transcription. Each column is annotated with what
its own values suggest - all "Normal", unique per row, 3 values - which is
usually enough to spot the health column at a glance.
Notes:
- Numeric OIDs only. A MIB name can't be resolved without its MIB file,
so
sysDescr.0is refused before anything touches the network. - Reading a table is capped, and a truncated result says so - go one level deeper rather than trusting a partial view.
- Unreachable device, wrong community, or no applicable profile come back as a message in the dialog, not a failed request. Small BMCs sometimes time out when browsed several times in quick succession; that reads as an SNMP error, never as "nothing there", so a retry is the obvious next move.
- Standard tables are offered in the field too -
hrDeviceTable,hrStorageTable,entPhySensorTable,entPhysicalTable- and are worth trying before the vendor tree, since they mean the same thing on every agent.
Exploring reads from the device and writes nothing, but it does make the server query arbitrary operator-supplied OIDs on that host, so it takes the same device change permission as the rest of the SNMP tooling.
Where sensors live¶
A sensor is a property of the hardware model, not of one box: an OID that reads drive health on one chassis reads it on every one you own. So there are three places to work with them, all editing the same records:
| Where | For |
|---|---|
| Device → SNMP tab | Explore a live device's OIDs, define a sensor from what you find, poll it, read the last values. |
| Device type → Sensors tab | The definitions this model carries. Every device of the type inherits them. |
| Settings → SNMP sensors | The whole catalog: search, duplicate, delete, and export/import packs. |
A sensor bound to a device type applies only to that model; one left unbound applies to all types. The device-type tab deliberately lists only its own, so you can't edit a shared definition by accident while looking at one model.
Sharing sensors as a pack¶
Working out that a Lenovo chassis reports drive health at
1.3.6.1.4.1.2.3.51.3.1.12.2.1.3 is real work, and it's the same answer for
everyone with that chassis. Settings → SNMP sensors → Export pack writes the
tenant's sensors to a JSON file; Import pack reads one back.
A pack contains only definitions - OID, walk/scalar, value map, naming rule, apply mode. No credentials: sensors poll with the device's own SNMP profile, so there is nothing secret to leak.
- Device types travel as their name, not their id (ids are per-deployment). A sensor naming a type you don't have is still imported, just unbound - the import tells you which, so you can bind it in one click.
- Sensors are matched by slug. Re-importing updates in place instead of
piling up duplicates, and by default an existing slug is skipped so an
import can't quietly rewrite a sensor someone tuned. Tick Overwrite to
replace (that needs
changeaccess, not justadd). - The envelope is versioned (
danbyte_snmp_sensor_pack: 1), so a file from a future Danbyte is rejected with a clear message rather than half-applied.
API: GET /api/monitoring/snmp-sensors/export/ and
POST /api/monitoring/snmp-sensors/import/?replace=0|1.
Defining a sensor by hand¶
On a device's SNMP tab → Custom SNMP sensors card (or the device type's Sensors tab), Add sensor:
- OID - the numeric OID. A walk reads a table column (one value per component, e.g. per drive); a scalar reads one value.
- Reading is - which hardware kind these readings describe (disk, PSU…).
- Item name template - how each reading names/matches its
inventory item:
{index}is the walk row,{kind}the kind (e.g.Disk {index}→Disk 1,Disk 2). - Value → status - map each raw SNMP value to a status slug, e.g.
3 → active,4 → failed. Unmapped values leave the item untouched. - Never reported - status for parts the sensor covers that the agent never lists. See empty bays.
- Scope - limit the sensor to this device type, or apply it to all types (define once, reuse across every server of that model).
Poll sensors runs every applicable sensor with the device's SNMP profile and records what it read. What that does to your data depends on one setting:
A reading is observed data¶
Danbyte is a source of truth with drift visualisation, and a health reading is observed data like any other. So by default a sensor never writes: the reading is stored, and where it disagrees with the status you set, that difference is listed as drift for you to accept - exactly how interfaces behave.
That means a part carries both states, and you can see them at once:
- The Hardware tab shows the part's status with an amber drift pill
beside it; the pill's popover reads
Active → failed (Drive health: Critical) - set status, observed status, and the raw value behind it.
- The photo faceplate and 3D rack keep drawing the part in its set status and ring a drifting bay in amber, so the picture stays the source of truth and the disagreement is a separate signal. The bay's popover names what SNMP said.
- The drift inbox on the Monitoring tab lists it with the usual accept / dismiss, and a part the agent reports that Danbyte has no record of appears as a new part to accept rather than being created behind your back.
Accepting is the only thing that writes.
Opt in to automatic application
Tick Apply readings automatically on the sensor and it writes straight through instead - a failing disk turns red with nobody watching. Off by default, deliberately: that mode can overwrite a status a human set, so it has to be asked for. Flips are journaled either way.
The Redfish collector still writes directly
BMC/Redfish health has not been moved onto this path yet - it reconciles part statuses on every poll, as sensors used to. Sensors are the SoT-compliant path today.
How a reading finds its part¶
By name, and only by name. The template renders one name per reading and
that string must equal the inventory item's name exactly - disk{index} on a
walk whose rows are 0…6 produces disk0 … disk6, which matches parts called
exactly that.
There is no stored link here. Unlike an interface, which records the SNMP name the agent uses for it, a part carries nothing that says "this reading is mine". Two consequences worth knowing before you rename anything:
- Rename a part and the sensor stops finding it. It doesn't error - it creates a second item under the templated name, and the renamed one keeps its last status forever. Change the template alongside the name, or don't rename monitored parts.
- The template has to match how the agent indexes, not how you'd label the
bay. A 0-based walk with a
Disk {index}template producesDisk 0, so parts namedDisk 1 … Disk nwill all be missed and duplicated. Use Explore OIDs to see the real row indexes first.
The safest order is: read the table, note its indexes, then write a template that lands on the names you already have.
Empty bays¶
A device type's inventory templates stamp every bay a chassis has - 16 disk bays on a 16-bay server - while the agent only reports the bays that are populated. Without help, the nine empty bays on a 7-disk machine keep claiming to hold healthy hardware, on the Hardware tab and on the faceplate.
Set Never reported on the sensor to a status like Empty, and after each poll any part the sensor covers (same kind, same scope) that the agent didn't list flips to it. A bay that later gets a disk is picked up and marked healthy again on the next poll.
Only ever applied after a poll that returned something
An agent that answers with nothing - blocked column, wrong community, a subtree that moved - looks exactly like "every bay is empty". Acting on that would mark real disks missing, so a poll with no readings, or one that errored, changes nothing. Silence is never evidence.
Only the sensor's own kind is touched: a disk sensor can't mark the power supplies empty.
To find your vendor's OIDs, walk the BMC's enterprise tree
(snmpwalk -v2c -c <community> <bmc> 1.3.6.1.4.1) and consult its MIB - the
disk-status column is what you point the sensor at.
Permissions¶
- Read (view observed facts, view drift, view topology, learned MACs per device) - view on the device, within its site scope. A port's MAC list needs view on the interface; the network-wide learned list needs view on MAC addresses and lists only devices you may view.
- Poll now and Refresh MACs -
device.change. Refresh MACs checks it again when the job runs. - Accept drift (reconcile observed → intended) - requires
device.change. This is the one source-of-truth write in the whole feature, so it's gated like editing the device itself, not merely tenant membership. - Uplink: Always / Never - an interface edit, so
interface.change. - Manage profiles & bindings - gated to users who can change the device / manage settings. The MAC-tracking settings are tenant admin settings.
Three more per-tenant policies (Monitoring settings → SNMP discovery, all off by default) shape how SNMP meets your source of truth:
- Only update existing interfaces - SNMP never adds ports; drift and sync only touch fields (MAC, speed, VLAN, enabled) on interfaces you created.
- Skip unrouted VLAN pseudo-interfaces - Cisco lists every L2 VLAN in the
interface table (
unrouted VLAN 401); with this on they're ignored as the VLANs they are. Routed SVIs are unaffected. - Interface MAC from the MAC table - the MAC drift/sync value becomes the address learned on the port (the attached device, per the switch's MAC table) instead of the port's own hardware MAC. Ports with several learned MACs are left alone. Since 0.17 it reads the port's present sightings, so the switch's own addresses never count as a learner.
Port access-VLANs resolve against your existing VLANs by VLAN ID - ungrouped first, then grouped (virt-sync groups excluded) - before a new ungrouped VLAN is minted.
Where discovered IPs land (VRF): an interface's own VRF always wins. When it has none, a default VRF resolves most-specific first - device → device role → device type → site → the tenant default in Monitoring settings. Bind it where it fits: the Discovered-IP VRF select on the device's SNMP card, the device type's Monitoring card, or the site form's Monitoring section. With a policy bound, only prefixes in that VRF are candidates - no containing prefix there means the address is skipped rather than dropped into the wrong table.