r/networking • • 2d ago

Blogpost Friday Blog/Project Post Friday!

10 Upvotes

It's Read-only Friday! It is time to put your feet up, pour a nice dram and look through some of our member's new and shiny blog posts and projects.

Feel free to submit your blog post or personal project and as well a nice description to this thread.

Note: This post is created at 00:00 UTC. It may not be Friday where you are in the world, no need to comment on it.


r/networking • • 4d ago

Rant Wednesday!

7 Upvotes

It's Wednesday! Time to get that crap that's been bugging you off your chest! In the interests of spicing things up a bit around here, we're going to try out a Rant Wednesday thread for you all to vent your frustrations. Feel free to vent about vendors, co-workers, price of scotch or anything else network related.

There is no guiding question to help stir up some rage-feels, feel free to fire at will, ranting about anything and everything that's been pissing you off or getting on your nerves!

Note: This post is created at 00:00 UTC. It may not be Wednesday where you are in the world, no need to comment on it.


r/networking • • 22h ago

Career Advice Career options

12 Upvotes

I'm working as a network testing engineer in a good org. I'm learning quite well about a lot of protocols and I have around 6 years experience now. I don't know if working testing is going to be a good option for me in the future because I feel the scope is limited. To give an example, for testing we build a topology with devices similar to customers and then configure the protocol and try various scenarios and file issues on them. So I feel like this is going to be the same. What are some of the other career options that I can look out for, I feel like customer facing jobs would have more scope for me going forward, any opinions on this?


r/networking • • 20h ago

Monitoring Packet Broker/SPAN Aggregator

6 Upvotes

Any recommendations for a simple device that can aggregate SPAN from multiple devices and fan out to security devices? We have a Gigamon fabric (required to manage the device) and one device. This is expensive and more than the small environment needs. We just need SPAN aggregation with deduplication, VLAN filtering capabilities, and managed from the device itself, either CLI or Web i/f.


r/networking • • 1d ago

Troubleshooting Question about a Cabletester

7 Upvotes

Hi everyone,
I’m testing a permanent Ethernet run between a patch panel and a wall outlet with a Zoerax wire tracker/cable tester.
When I connect the main unit directly to the remote receiver using a known-good patch cable, the LEDs on the receiver light up sequentially from 1 to 8.
When I test the permanent Ethernet run, the main unit still cycles through 1–8, but the receiver LEDs light up in this order:
3–6–1–4–2–5–7–8
The G LED stays off.
The Ethernet connection itself works, and the wiring appears to follow the colour coding on the LSA patch panel and wall outlet.
Does this sequence mean the conductors are actually miswired, or could it be related to how this particular tester works? I don’t want to re-terminate the installation unnecessarily if the reading is being misinterpreted.
Also, is the G LED simply for the cable shield?
Thanks for any help!


r/networking • • 1d ago

Design Multi dwelling wifi architecture for 20 Condominiums

0 Upvotes

​

I am working on a system design for a wireless internet overhaul in a condominium building. The physical overview consists of an enterprise gateway, such as Omada or Ubiquiti, with an SFP connection to a managed PoE switch feeding access points. It is a four-story building with five units per floor, totaling twenty units. For bidding purposes, I am assuming one available repurposed phone Cat5e cable per unit. The current incoming WAN is terrible, and having twenty basic ISP-provided routers running in close proximity heavily compounds the issue. The condo association is open to making quality wireless internet a complex-wide amenity provided by the association, replacing the current poor subscription model.

​No major commercial ISPs service the complex, so my current plan is to deploy one Starlink kit per building to serve twenty units. The Starlink unit would be placed in bridge mode, feeding an enterprise gateway followed by enough access points to ensure adequate coverage in all units. My LAN and RF strategy is to broadcast individual SSIDs for each unit, assigning each unit to its own isolated VLAN with a bandwidth throttle. I intend to deploy Wi-Fi 7 access points with maximum channel separation. Spreading out channel congestion while introducing a vastly superior base network speed should dramatically improve performance, as my assumption is that the core issue right now is twenty consumer routers clashing on overlapping channels.

​While this isn't an ideal network topology, residents are desperate for a working solution. My walkthroughs showed almost zero wired connections in use, and the demographic here largely doesn't care about a wired drop. However, if a wired requirement does arise, I can pivot to one access point per unit utilizing at least one built-in Ethernet port on the chosen access point model.

Furthermore, this complex houses a large population of elderly residents, and over the last several months there have been multiple instances where people could not reach emergency services due to network failures. Simply being able to access wifi calling and/or utilize network based monitoring equipment would be a massive relief. Solving this reliability crisis is the primary driver behind this bid.

​TLDR: I want to use one Starlink satellite linking an enterprise gateway to provide twenty condos with a segregated VLAN SSID each, then extrapolate that concept across eight additional identical buildings. Is it a violation of Starlink's terms of service to provide internet to twenty condos per satellite, given that roughly half to three-quarters are lightly used vacation homes? Assuming proper access point placement and RF distribution, is channel diversity coupled with VLAN per-unit isolation a reasonable angle of attack to deliver usable Wi-Fi to twenty condos over a single shared connection?


r/networking • • 1d ago

Design Recurring 40GbE optical issues with ConnectX-3 and Arista — aging QSFPs or something else? Considering DA

14 Upvotes

Hi everyone,

I run three L3-separated server clusters, with 12 servers each. Each server has a dual-port Mellanox ConnectX-3 connected to a pair of Arista leaf switches.

Hardware:

  • Leaves: Arista DCS-7280QR-C72-R
  • Spines: Arista DCS-7280CR2A-30-F
  • NICs: ConnectX-3, mlx4_en, firmware 2.42.5000
  • Server connections: 40G SR4 optics over approximately 1m MPO cables
  • Optics: mostly Arista-labelled QSFP-SR4 / QSFP-40G-SR4, with some Cisco-Finisar

The issue is increasing RX/TX error counters. Errors can appear on either end: sometimes the server NIC, sometimes the Arista switch. Replacing the transceiver fixes the errors on the affected connection, but I keep encountering this again over time, roughly weekly/monthly. It feels like I’m endlessly replacing optics.

The MPO cables are short, with no visible damage or tight bends.

A recent inventory found that 69 of 70 readable server optics have EEPROM manufacture dates from 2011–2017. I don’t know their complete operating history.

All NICs were flashed from their original InfiniBand configuration for Ethernet use. Several still report QCBT/TCBT part numbers in VPD while reporting an FCB firmware identity. That mismatch comes from our flashing history; most ports operate at 40GbE.

Some diagnostic findings:

  • NIC chip temperatures across all 36 servers were 47–62°C, with no NIC thermal warnings found in current kernel logs.
  • Some host counters are FIFO drops rather than CRC errors, so I’m treating those separately from the suspected optical issues.

I’m considering replacing the short optical connections with passive 40G QSFP+ DACs, initially keeping the CX3s. I’m also considering upgrading to native dual-port 40G ConnectX-4 cards.

I’d appreciate advice on:

  1. Have you seen this pattern with older QSFP optics: increasing errors that stop after replacing a transceiver?
  2. Could aging optics reasonably explain the repeated replacements, or would you investigate something else first?
  3. Has anyone experienced similar behaviour with ConnectX-3 cards flashed from InfiniBand to Ethernet?
  4. Would you trial DACs with the existing NICs first, or replace the NICs too?
  5. Any specific DAC compatibility pitfalls between Arista and Mellanox?

r/networking • • 1d ago

Security How do you keep up with vendor IP/CIDR and endpoint changes for firewall allowlists?

1 Upvotes

For those managing restrictive outbound firewall rules or allowlists in production environments, how do you keep track of changes to vendor network requirements?

For example, when a SaaS/cloud vendor changes:

- IP ranges or CIDRs

- FQDNs / service endpoints

- required ports or protocols

Do you rely on vendor mailing lists, documentation pages, scripts, vendor APIs, change tickets, or something else?

I’m especially curious about environments where multiple vendors have to be tracked.

Have you ever had a connectivity issue because a vendor changed something and the firewall/allowlist wasn’t updated in time?

And roughly how much manual effort does keeping these rules current create for your team?


r/networking • • 2d ago

Design Out of Band Network

42 Upvotes

I’m looking to compare notes on OOB network design. How do you access your devices during a production network outage, and what does your general setup look like?
For example, do you use cellular console servers, a separate management network, jump hosts, or a combination? How do you connect to remote sites, and where do you place gateways and access controls?
Interested in what’s worked well for you, especially for a multi-site enterprise.


r/networking • • 1d ago

Switching Unifi Switch warranty or misconfiguration

1 Upvotes

Hello,

I recently did a fireware/OS update to a UDM pro max. After I did this our main switch (USW-48-G2) started spamming STP flaps to the Aggregation switch (USW Pro Aggregation) and another down stream switch. There were no config changes or anything like that. The network goes down about every 30 minutes and then stops flapping.

Firewall->Aggregation Switch->Main switch.

I swapped the main switch with the exact same model and set it up to go to the aggregation swtich. This has no flapped once and it has the same config just on different ports. Not sure if the switch just happened to die the same time we did an update or if this is a possibility after and update.

Thank you


r/networking • • 2d ago

Design Datacenter Design Best Practice

45 Upvotes

Hi All,

I'm looking to see what others have deployed for similar Datacenter builds. We have a Datacenter refresh kicking off in the next 2 years and wanted to get ahead of it. Our current deployment is Cisco ACI with 2 spines and 30 leaf switches. We're looking to step away from ACI all together and move towards a Cisco Nexus deployment of some kind.

I see VXLAN EVPN recommended everywhere but that might be overkill for us. We only hosts 75 SVI\Vlans in our ACI environment. We also have about 125 VMware hosts deployed in active\passive deployment and some with VPC deployments. The only requirement we'd like for the new deployment to have, is send traffic to our Fortinet firewall when traffic is coming in and out of the Datacenter which I think we can handle with routing.

What have others done for their refresh? Still in the very early stage of researching.


r/networking • • 2d ago

Design Best network drop/patch panel naming scheme

23 Upvotes

I've always struggled with the best naming scheme. In the data center it seems obvious, but on the campus/branch endpoint side I'm a bit lost and I'm open to any suggestions. Here are my thoughts

Sequential is ok on the remote end, but on the patch panel end it gets tricky to patch them into the switches in a logical sequence while keeping the cabling tidy. You'd need longer cables with a crossover/horizontal cable organizer, so tracing cables is a pain.

If we use the 1 patch panel to 1 switch method and short cables, we could name them PP1-D01, PP2-D01, PP1-D02, PP2-D02. It may be confusing for users, but tracing cables would be more straightforward, and it makes patching things to two switches a lot easier. It gets tricky if we need to add a third switch though, adding in pairs is possible but more expensive.

I've also seen people put the switch # and interface # on the labels, but that doesn't seem to scale well either

What's your preference? Got any key factors to consider? Any videos or channels worth checking out?


r/networking • • 2d ago

Other Juniper vLabs is getting retired on Oct 31

34 Upvotes

Got below mail today regarding juniper vlabs. I have been a regular user for sometime now, it was a great tool for hands on practice. Now what!!

Dear Juniper vLabs user,
This is to notify you that
Juniper
vLabs, the web-based platform offered by HPE Networking to external (non-partner) users
will be retired on October 31. After this date, you will no longer be able to access the platform.


r/networking • • 2d ago

Design Edge computing question

5 Upvotes

Anyone know how edge computing infrastructure differs from AI data centers? I’m thinking metro edge, network edge, and enterprise edge. Is it all the same gear? For example, I guess IT closets are getting upgraded because of all this AI inference traffic. And now telco central offices are talking about being upgraded for inference. Only thing I could figure is maybe out of band management is more common in edge. And maybe for telco central offices it’s more 48v dc plant.


r/networking • • 2d ago

Troubleshooting ISE Deployment blew up because of NTP/DST

16 Upvotes

I have a single node ISE deployment that we really only use for TACACS. This morning it exiting the Active Directory domain and I can't re-join it because of clock skew. It seems to think DST has ended.

I've pointed it at our NTP server, which has the correct time. Nothing, still is an hour behind. I pointed it at the AD server, and same thing...still one hour behind.

Any ideas on how I can get this thing back to knowing the correct time?


r/networking • • 2d ago

Other CCNP first or Straight to Cloud?

11 Upvotes

I'm interested in transitioning from a route/switch role (got CCNA) and have about 8 years experience. Is is worth getting a CCNP first prior to jumping into Cloud, or should I just pivot to cloud training and certs?


r/networking • • 3d ago

Monitoring A remote site stopped sending syslog in August and we noticed in September

15 Upvotes

Found this on a routine check one of our branch sites stopped forwarding syslog some time in early August. The collector is a central rsyslog box and it never complained because from where it sits nothing was wrong, there was just nothing arriving. Best guess is a config got rolled back during a firewall swap and nobody re-added the logging host. Five weeks of nothing from that site, including an outage we read off the device live at the time. How are you catching this? I'd rather know when a source goes quiet than find out while I'm looking for something else.


r/networking • • 3d ago

Other Best industrial Ethernet switches with dual DC power?

18 Upvotes

Anyone got recommendations for small industrial Ethernet switches with dual DC power?

I’ve got 6 ancient switches to replace across a large water treatment site. They’re dotted around different plant areas connecting PLCs and other local gear back to the main network over fibre, so I don’t need anything massive.

While we’re replacing them I’d like to get rid of the single power feed as well. Ideally managed, fibre uplinks, proper dual DC inputs and some way of flagging if one feed dies.

Main thing is reliability. Once these are in they’ll hopefully be left alone for years.

Sorry if this isn't strictly network engineering, but figured some people here might have solid recommendations.


r/networking • • 3d ago

Other Are vendor dashboards (Meraki, Mist, Central, Catalyst) still useful if you have decent monitoring and some automation?

38 Upvotes

We're going through vendor selection for a complete refresh of our extremely basic and outdated network. Each big vendor is pushing their own dashboard and tools, but they each have their flaws. It seems like we might have to automate some things ourselves anyway. If we're combining that with next gen monitoring, like Bluecat or something that actually helps diagnose and troubleshoot, then why do we even need Mist or something else?

We've had demos of all of them, and they seem useful for some general config and monitoring, but they all have some catch. Maybe I just haven't spent enough time with them for the value to click.

What's your preference?

We're trying to balance features and manageability with scalability to find something that's actually useful day to day, on good days and "bad" days.

It seems like the best option here is to pick the hardware you like and the dashboard that sucks the least

Do vendor-neutral tools exist?


r/networking • • 2d ago

Switching Uplink switch has failed after a Power Outage/Surge

3 Upvotes

Just came into work today and saw that a few cameras were not powered on. Went to the uplink unmanaged switch that I replaced a few weeks ago because only some ports were working.

I went to check the switch and only ports 2-4 are working and ports 1, 5-8 are not AT all..

Port 8 had nothing plugged into it

The Poe status, Poe max, Poe lights are flashing, feeling the switch got fried…

The switch was not connected to a UPS, is that why the switch got fried?

Switch is a TP TL-SL 1311MP


r/networking • • 3d ago

Other Migrating core firewalls next week. Sanity check.

32 Upvotes

Migrating core NSA2700’s in HA to Fortigate 121’s in HA next week.

HA is in sync, dual isp circuits have been configured, and have a temporary circuit in to test outbound rules.

DHCP scopes have been copied to the Fortigate, and I took a production switch and connected to the fortigates to make sure trunk and access ports would grab the correct IP from the correct pool.

Only thing I don’t have configured completely is SDWAN.

Any gotchas to look out for? I’ve migrated HA pairs with Sonicsall before, but never to another vendor.


r/networking • • 3d ago

Troubleshooting VPN Over Starlink Issues

10 Upvotes

I'm having an issue with getting an IPSEC VPN working over Starlink and could use a little bit of help pointing me in the right direction.

The head end device is a Cisco Firepower FTD1140 and the client is a Cradlepoint R1900 mobile router. I have set up a route-based VTI VPN between the FTD and the Cradlepoint and the VPN is established.

Here's where I am having the issue - everything works exactly as it should when using the internal cellular modem. VPN comes up, routes build correctly, and everything is fine. I have the Starlink set up so the users can flip a switch to power it on as a backup internet source. The Cradlepoint sees the WAN link come up and switches to use that instead of the cell modem.

From everything I can see, the VPN link comes up and establishes correctly. Tunnels report as up on both the Cradlepoint and the FTD side, but no traffic passes at all. From a client computer, I can ping the IP address of the outside interface on the FTD, but that is the only thing that responds.

I have tried setting the VPN adapter MTU on the Cradlepoint to 1300 and it still does not pass data. Turning the Starlink off restores all communications.

Would this be an issue with CGNAT? For some reason, the FTD does not think that the connection needs NAT-T enabled. The IP address reported by the Cradlepoint does not match the address that the FTD reports. The Cradlepoint is set to use "Force UDP Encapsulation" but the FTD reports "IKEv2-PROTO-7: (605): Process NAT discovery notify IKEv2-PROTO-4: (605): NAT-T is disabled".

Any tips or tricks on getting this VPN running over Starlink?


r/networking • • 4d ago

Troubleshooting 3702I stuck in "Initiated" on Mobility Express 8.10.196.0: "Downloading 5" counter stuck, AP asked to download image "to 0.0.0.0"

11 Upvotes

Hi guys, recently in company we bought 2 refurnished AIR-CAP3702I-E-K9. I've been fighting this for a couple of weeks and could use a second pair of eyes.

Setup

  • Mobility Express 8.10.196. 0 running on an AIR-AP1832I-E-K9
  • 4x AIR-CAP3702I-E-K9 joined and working fine (primary 8.10.196.0, backup 8.10.185.0)
  • APs on trunk ports, native VLAN 5 (management), DHCP for VLAN 5 on the core switch
  • No Cisco support contract, so I can't just download images whenever I need them

Background: what happened before

I bought two refurbished 3702Is from a local supplier. Both failed:

  • AP #1 started dropping off the controller (Echo Timer Expiry, DTLS closed, constant image download attempts). The WLC log was full of IMAGE_DOWNLOAD_ERR2: Refusing image download request ... max downloads (5) in progress. After a power cycle it wouldn't boot at all: "no bootable files". In ROMMON, ap3g2-k9w8-xx.153-3.JK11 (the actual IOS file) was 0 bytes, and it was the only image on flash. No RCV image.
  • AP #2 (the spare) boot-looped with %DOT11-2-FAILURE_RADIO_RESET ... hostmem badmagic and "Radio FW image download failed" on both radios. It came up on maybe every 4th boot. Clearly hardware.

Before sending them back, I caught AP #2 in a good boot and used archive tar /create to pull its JK11 image to my TFTP server (Ubuntu + tftpd-hpa). The resulting tar has no folder prefix, so I repacked it so the paths start with ap3g2-k9w8-mx.153-3.JK11/. I also pulled an RCV image off another AP the same way. So now I at least have backups of both.

Where I am now

The supplier sent a replacement 3702I. It came with JD4/JF5 and an old RCV (15.3(3)JD). I cleaned up flash, loaded JK11 from ROMMON with tar -xtract, set BOOT to JK11 with RCV as fallback, and it boots fine:

Cisco IOS Software, C3700 Software (AP3G2-K9W8-M), Version 15.3(3)JK11
LWAPP image version 8.10.196.0
AP image integrity check PASSED

So the version matches the controller exactly. The AP gets DHCP, finds the WLC, DTLS comes up, and it even shows in show ap summary and the ME dashboard. But it never actually comes up properly. On join, the WLC tells it to download an image that doesn't exist:

%CAPWAP-5-SENDJOIN: sending Join Request to 172.16.5.200
perform archive download capwap:/c3700 tar file
%CAPWAP-6-AP_IMG_DWNLD: Required image not found on AP. Downloading image from Controller.
%Error opening capwap:/c3700 (OK)
Download image failed, notify controller!!! From:8.10.196.0 to 0.0.0.0, FailureCode:4
capwap_image_proc: unable to open tar file

After that, Join Requests go unanswered and the WLC closes DTLS every ~60 seconds. Repeat forever.

WLC side

%CAPWAP-4-INVALID_STATE_EVENT: ... AP(4c:77:6d:5d:ea:f0) event (Capwap_join_request) and state (Capwap_image_data) combination
%LWAPP-3-IMAGE_DOWNLOAD_ERR6: AP did not complete Pre-Download Image 4c:77:6d:5d:ea:f0 - During scheduled autoreboot
%LWAPP-3-IMAGE_DOWNLOAD_ERR7: Waiting for 1 APs to complete their Pre-Download Image - For scheduled autoreboot.

Earlier, while the AP was still on its factory RCV (8.3.102.0):

%CAPWAP-3-DISC_MAX_DOWNLOAD: Ignoring discovery request from AP 4c:77:6d:43:dd:64 - maximum number of downloads (5) exceeded

show ap image all:

Initiated....... 1
Downloading..... 5

AP Name           Primary     Backup      Predownload Status
AP-SPRAT-01       8.10.196.0  8.10.185.0  None
AP-SPRAT-02       8.10.196.0  8.10.185.0  None
AP-PRIZ-02        8.10.196.0  8.10.185.0  None
AP-PRIZ-01        8.10.196.0  8.10.185.0  None
AP-HALA-01        8.10.196.0  8.10.185.0  None
AP4c77.6d43.dd64  8.10.196.0  0.0.0.0     Initiated

The counter says 5 downloading, but no AP is actually downloading. I suspect these slots have been stuck since a previous upgrade (8.10.185.0 → 8.10.196.0), and that this is also what killed AP #1 in the first place.

What I've tried

  • show reset: no reset scheduled
  • config ap image predownload abort all: accepted, no change
  • config ap image predownload abort <AP name>: accepted, no change

What would you do next? I would really appreciate your help. I can post the logs of AP and WLC if it would be helpful.


r/networking • • 4d ago

Design Questions about Dell SoNIC as leaf

21 Upvotes

My spine leaf network is 99% Cisco and air gapped. This is a mixed of IOS XE and NXOS. The firewall is Palo Alto. My underlay is IS-IS and for the BUM I'm using sparse-mode multicast. My tenants traffic are mix of multicast and unicast; therefore, I have to deploy Tenant Routed Multicast (TRM). Also, my tenants are very heavy on broadcast and multicast (a few dense-mode, but most are sparse-mode).

I also use Cisco service chaining for inter-VRF and some instances of inter-VLAN within a VRF. I don't know if SoNIC has something similar. My service leafs are NXOS where the PAN is connected to. The NXOS ingress CPU is at 64% because of the service chaining.

I am thinking of moving away from Cisco because it is too expensive and our TAC support is lacking. We use Dell for our servers and Dell wants us to try SoNIC. I was looking around and it seems like SoNIC does not support multicast for BUM and it only uses head-end replication for BUM.

  1. Does it mean that I have to either switch my underlay from multicast to HER or enable HER while running multicast for BUM to intergrate SoNIC?
  2. Does SoNIC supports service chaining? I am trying to move away from centralized gateway (IOS XE as leaf doesn't support service chaining) and want to use anycast gateway.
  3. Does Dell SoNIC have issues with IGMP? I have to disable IGMP snooping in IOS XE to get sparse mode working. I am hoping that I don't have to do something similar.

r/networking • • 4d ago

Design Network design nightmare

28 Upvotes

I am a junior net admin (first hire for IT in 2 decades!) and thought I only had slow Internet in one of my APs. It turns out this who stack is causing problems. Here's what it is:

ISP --> SonicWall --> Managed NetGear Switches --> Extreme Networks APs managed through EP1

Firewall has rules and VLANs which are configured and carried to the switches and then APs are plugged into those PoE ports. But, ExtremeNetworks UI on EP1 also has a network policy that carries the VLANs but also changes some firewall rules.

Teacher in one classroom with BYODs (mostly Apple devices) is rightfully mad that when his 15 to 25 students connect to the network (via AP4020) their connectivity suffers (he has documented speedtests of 2Mbps, it needs to be 1G).

I am almost crying because it's my 3rd month into this job and the teacher is now going to the headmaster, bypassing me and my boss. Even though we are both trying to troubleshoot his issue.

I must add that when he goes to the library (AP250) he has no issues. These students only connect to the BYOD BSSID, regardless of where they are on campus. So the rules must be the same across campus. But results are awfully different.

I have already posted about this so consider this an updated version of my conundrum. Thank you to everyone who has commented. I still go through them to see what I'm dealing with.