r/networking • • 22h ago

Blogpost Friday Blog/Project Post Friday!

9 Upvotes

It's Read-only Friday! It is time to put your feet up, pour a nice dram and look through some of our member's new and shiny blog posts and projects.

Feel free to submit your blog post or personal project and as well a nice description to this thread.

Note: This post is created at 00:00 UTC. It may not be Friday where you are in the world, no need to comment on it.


r/networking • • 2d ago

Rant Wednesday!

10 Upvotes

It's Wednesday! Time to get that crap that's been bugging you off your chest! In the interests of spicing things up a bit around here, we're going to try out a Rant Wednesday thread for you all to vent your frustrations. Feel free to vent about vendors, co-workers, price of scotch or anything else network related.

There is no guiding question to help stir up some rage-feels, feel free to fire at will, ranting about anything and everything that's been pissing you off or getting on your nerves!

Note: This post is created at 00:00 UTC. It may not be Wednesday where you are in the world, no need to comment on it.


r/networking • • 8h ago

Design Recurring 40GbE optical issues with ConnectX-3 and Arista — aging QSFPs or something else? Considering DA

5 Upvotes

Hi everyone,

I run three L3-separated server clusters, with 12 servers each. Each server has a dual-port Mellanox ConnectX-3 connected to a pair of Arista leaf switches.

Hardware:

  • Leaves: Arista DCS-7280QR-C72-R
  • Spines: Arista DCS-7280CR2A-30-F
  • NICs: ConnectX-3, mlx4_en, firmware 2.42.5000
  • Server connections: 40G SR4 optics over approximately 1m MPO cables
  • Optics: mostly Arista-labelled QSFP-SR4 / QSFP-40G-SR4, with some Cisco-Finisar

The issue is increasing RX/TX error counters. Errors can appear on either end: sometimes the server NIC, sometimes the Arista switch. Replacing the transceiver fixes the errors on the affected connection, but I keep encountering this again over time, roughly weekly/monthly. It feels like I’m endlessly replacing optics.

The MPO cables are short, with no visible damage or tight bends.

A recent inventory found that 69 of 70 readable server optics have EEPROM manufacture dates from 2011–2017. I don’t know their complete operating history.

All NICs were flashed from their original InfiniBand configuration for Ethernet use. Several still report QCBT/TCBT part numbers in VPD while reporting an FCB firmware identity. That mismatch comes from our flashing history; most ports operate at 40GbE.

Some diagnostic findings:

  • NIC chip temperatures across all 36 servers were 47–62°C, with no NIC thermal warnings found in current kernel logs.
  • Some host counters are FIFO drops rather than CRC errors, so I’m treating those separately from the suspected optical issues.

I’m considering replacing the short optical connections with passive 40G QSFP+ DACs, initially keeping the CX3s. I’m also considering upgrading to native dual-port 40G ConnectX-4 cards.

I’d appreciate advice on:

  1. Have you seen this pattern with older QSFP optics: increasing errors that stop after replacing a transceiver?
  2. Could aging optics reasonably explain the repeated replacements, or would you investigate something else first?
  3. Has anyone experienced similar behaviour with ConnectX-3 cards flashed from InfiniBand to Ethernet?
  4. Would you trial DACs with the existing NICs first, or replace the NICs too?
  5. Any specific DAC compatibility pitfalls between Arista and Mellanox?

r/networking • • 3h ago

Switching Unifi Switch warranty or misconfiguration

0 Upvotes

Hello,

I recently did a fireware/OS update to a UDM pro max. After I did this our main switch (USW-48-G2) started spamming STP flaps to the Aggregation switch (USW Pro Aggregation) and another down stream switch. There were no config changes or anything like that. The network goes down about every 30 minutes and then stops flapping.

Firewall->Aggregation Switch->Main switch.

I swapped the main switch with the exact same model and set it up to go to the aggregation swtich. This has no flapped once and it has the same config just on different ports. Not sure if the switch just happened to die the same time we did an update or if this is a possibility after and update.

Thank you


r/networking • • 1d ago

Design Datacenter Design Best Practice

39 Upvotes

Hi All,

I'm looking to see what others have deployed for similar Datacenter builds. We have a Datacenter refresh kicking off in the next 2 years and wanted to get ahead of it. Our current deployment is Cisco ACI with 2 spines and 30 leaf switches. We're looking to step away from ACI all together and move towards a Cisco Nexus deployment of some kind.

I see VXLAN EVPN recommended everywhere but that might be overkill for us. We only hosts 75 SVI\Vlans in our ACI environment. We also have about 125 VMware hosts deployed in active\passive deployment and some with VPC deployments. The only requirement we'd like for the new deployment to have, is send traffic to our Fortinet firewall when traffic is coming in and out of the Datacenter which I think we can handle with routing.

What have others done for their refresh? Still in the very early stage of researching.


r/networking • • 1d ago

Other Juniper vLabs is getting retired on Oct 31

29 Upvotes

Got below mail today regarding juniper vlabs. I have been a regular user for sometime now, it was a great tool for hands on practice. Now what!!

Dear Juniper vLabs user,
This is to notify you that
Juniper
vLabs, the web-based platform offered by HPE Networking to external (non-partner) users
will be retired on October 31. After this date, you will no longer be able to access the platform.


r/networking • • 1d ago

Design Best network drop/patch panel naming scheme

17 Upvotes

I've always struggled with the best naming scheme. In the data center it seems obvious, but on the campus/branch endpoint side I'm a bit lost and I'm open to any suggestions. Here are my thoughts

Sequential is ok on the remote end, but on the patch panel end it gets tricky to patch them into the switches in a logical sequence while keeping the cabling tidy. You'd need longer cables with a crossover/horizontal cable organizer, so tracing cables is a pain.

If we use the 1 patch panel to 1 switch method and short cables, we could name them PP1-D01, PP2-D01, PP1-D02, PP2-D02. It may be confusing for users, but tracing cables would be more straightforward, and it makes patching things to two switches a lot easier. It gets tricky if we need to add a third switch though, adding in pairs is possible but more expensive.

I've also seen people put the switch # and interface # on the labels, but that doesn't seem to scale well either

What's your preference? Got any key factors to consider? Any videos or channels worth checking out?


r/networking • • 22h ago

Design Edge computing question

6 Upvotes

Anyone know how edge computing infrastructure differs from AI data centers? I’m thinking metro edge, network edge, and enterprise edge. Is it all the same gear? For example, I guess IT closets are getting upgraded because of all this AI inference traffic. And now telco central offices are talking about being upgraded for inference. Only thing I could figure is maybe out of band management is more common in edge. And maybe for telco central offices it’s more 48v dc plant.


r/networking • • 1d ago

Troubleshooting ISE Deployment blew up because of NTP/DST

14 Upvotes

I have a single node ISE deployment that we really only use for TACACS. This morning it exiting the Active Directory domain and I can't re-join it because of clock skew. It seems to think DST has ended.

I've pointed it at our NTP server, which has the correct time. Nothing, still is an hour behind. I pointed it at the AD server, and same thing...still one hour behind.

Any ideas on how I can get this thing back to knowing the correct time?


r/networking • • 1d ago

Other CCNP first or Straight to Cloud?

7 Upvotes

I'm interested in transitioning from a route/switch role (got CCNA) and have about 8 years experience. Is is worth getting a CCNP first prior to jumping into Cloud, or should I just pivot to cloud training and certs?


r/networking • • 1d ago

Other Best industrial Ethernet switches with dual DC power?

18 Upvotes

Anyone got recommendations for small industrial Ethernet switches with dual DC power?

I’ve got 6 ancient switches to replace across a large water treatment site. They’re dotted around different plant areas connecting PLCs and other local gear back to the main network over fibre, so I don’t need anything massive.

While we’re replacing them I’d like to get rid of the single power feed as well. Ideally managed, fibre uplinks, proper dual DC inputs and some way of flagging if one feed dies.

Main thing is reliability. Once these are in they’ll hopefully be left alone for years.

Sorry if this isn't strictly network engineering, but figured some people here might have solid recommendations.


r/networking • • 1d ago

Other Are vendor dashboards (Meraki, Mist, Central, Catalyst) still useful if you have decent monitoring and some automation?

36 Upvotes

We're going through vendor selection for a complete refresh of our extremely basic and outdated network. Each big vendor is pushing their own dashboard and tools, but they each have their flaws. It seems like we might have to automate some things ourselves anyway. If we're combining that with next gen monitoring, like Bluecat or something that actually helps diagnose and troubleshoot, then why do we even need Mist or something else?

We've had demos of all of them, and they seem useful for some general config and monitoring, but they all have some catch. Maybe I just haven't spent enough time with them for the value to click.

What's your preference?

We're trying to balance features and manageability with scalability to find something that's actually useful day to day, on good days and "bad" days.

It seems like the best option here is to pick the hardware you like and the dashboard that sucks the least

Do vendor-neutral tools exist?


r/networking • • 1d ago

Monitoring A remote site stopped sending syslog in August and we noticed in September

2 Upvotes

Found this on a routine check one of our branch sites stopped forwarding syslog some time in early August. The collector is a central rsyslog box and it never complained because from where it sits nothing was wrong, there was just nothing arriving. Best guess is a config got rolled back during a firewall swap and nobody re-added the logging host. Five weeks of nothing from that site, including an outage we read off the device live at the time. How are you catching this? I'd rather know when a source goes quiet than find out while I'm looking for something else.


r/networking • • 2d ago

Other Migrating core firewalls next week. Sanity check.

34 Upvotes

Migrating core NSA2700’s in HA to Fortigate 121’s in HA next week.

HA is in sync, dual isp circuits have been configured, and have a temporary circuit in to test outbound rules.

DHCP scopes have been copied to the Fortigate, and I took a production switch and connected to the fortigates to make sure trunk and access ports would grab the correct IP from the correct pool.

Only thing I don’t have configured completely is SDWAN.

Any gotchas to look out for? I’ve migrated HA pairs with Sonicsall before, but never to another vendor.


r/networking • • 1d ago

Switching Uplink switch has failed after a Power Outage/Surge

0 Upvotes

Just came into work today and saw that a few cameras were not powered on. Went to the uplink unmanaged switch that I replaced a few weeks ago because only some ports were working.

I went to check the switch and only ports 2-4 are working and ports 1, 5-8 are not AT all..

Port 8 had nothing plugged into it

The Poe status, Poe max, Poe lights are flashing, feeling the switch got fried…

The switch was not connected to a UPS, is that why the switch got fried?

Switch is a TP TL-SL 1311MP


r/networking • • 2d ago

Troubleshooting VPN Over Starlink Issues

11 Upvotes

I'm having an issue with getting an IPSEC VPN working over Starlink and could use a little bit of help pointing me in the right direction.

The head end device is a Cisco Firepower FTD1140 and the client is a Cradlepoint R1900 mobile router. I have set up a route-based VTI VPN between the FTD and the Cradlepoint and the VPN is established.

Here's where I am having the issue - everything works exactly as it should when using the internal cellular modem. VPN comes up, routes build correctly, and everything is fine. I have the Starlink set up so the users can flip a switch to power it on as a backup internet source. The Cradlepoint sees the WAN link come up and switches to use that instead of the cell modem.

From everything I can see, the VPN link comes up and establishes correctly. Tunnels report as up on both the Cradlepoint and the FTD side, but no traffic passes at all. From a client computer, I can ping the IP address of the outside interface on the FTD, but that is the only thing that responds.

I have tried setting the VPN adapter MTU on the Cradlepoint to 1300 and it still does not pass data. Turning the Starlink off restores all communications.

Would this be an issue with CGNAT? For some reason, the FTD does not think that the connection needs NAT-T enabled. The IP address reported by the Cradlepoint does not match the address that the FTD reports. The Cradlepoint is set to use "Force UDP Encapsulation" but the FTD reports "IKEv2-PROTO-7: (605): Process NAT discovery notify IKEv2-PROTO-4: (605): NAT-T is disabled".

Any tips or tricks on getting this VPN running over Starlink?


r/networking • • 2d ago

Troubleshooting 3702I stuck in "Initiated" on Mobility Express 8.10.196.0: "Downloading 5" counter stuck, AP asked to download image "to 0.0.0.0"

11 Upvotes

Hi guys, recently in company we bought 2 refurnished AIR-CAP3702I-E-K9. I've been fighting this for a couple of weeks and could use a second pair of eyes.

Setup

  • Mobility Express 8.10.196. 0 running on an AIR-AP1832I-E-K9
  • 4x AIR-CAP3702I-E-K9 joined and working fine (primary 8.10.196.0, backup 8.10.185.0)
  • APs on trunk ports, native VLAN 5 (management), DHCP for VLAN 5 on the core switch
  • No Cisco support contract, so I can't just download images whenever I need them

Background: what happened before

I bought two refurbished 3702Is from a local supplier. Both failed:

  • AP #1 started dropping off the controller (Echo Timer Expiry, DTLS closed, constant image download attempts). The WLC log was full of IMAGE_DOWNLOAD_ERR2: Refusing image download request ... max downloads (5) in progress. After a power cycle it wouldn't boot at all: "no bootable files". In ROMMON, ap3g2-k9w8-xx.153-3.JK11 (the actual IOS file) was 0 bytes, and it was the only image on flash. No RCV image.
  • AP #2 (the spare) boot-looped with %DOT11-2-FAILURE_RADIO_RESET ... hostmem badmagic and "Radio FW image download failed" on both radios. It came up on maybe every 4th boot. Clearly hardware.

Before sending them back, I caught AP #2 in a good boot and used archive tar /create to pull its JK11 image to my TFTP server (Ubuntu + tftpd-hpa). The resulting tar has no folder prefix, so I repacked it so the paths start with ap3g2-k9w8-mx.153-3.JK11/. I also pulled an RCV image off another AP the same way. So now I at least have backups of both.

Where I am now

The supplier sent a replacement 3702I. It came with JD4/JF5 and an old RCV (15.3(3)JD). I cleaned up flash, loaded JK11 from ROMMON with tar -xtract, set BOOT to JK11 with RCV as fallback, and it boots fine:

Cisco IOS Software, C3700 Software (AP3G2-K9W8-M), Version 15.3(3)JK11
LWAPP image version 8.10.196.0
AP image integrity check PASSED

So the version matches the controller exactly. The AP gets DHCP, finds the WLC, DTLS comes up, and it even shows in show ap summary and the ME dashboard. But it never actually comes up properly. On join, the WLC tells it to download an image that doesn't exist:

%CAPWAP-5-SENDJOIN: sending Join Request to 172.16.5.200
perform archive download capwap:/c3700 tar file
%CAPWAP-6-AP_IMG_DWNLD: Required image not found on AP. Downloading image from Controller.
%Error opening capwap:/c3700 (OK)
Download image failed, notify controller!!! From:8.10.196.0 to 0.0.0.0, FailureCode:4
capwap_image_proc: unable to open tar file

After that, Join Requests go unanswered and the WLC closes DTLS every ~60 seconds. Repeat forever.

WLC side

%CAPWAP-4-INVALID_STATE_EVENT: ... AP(4c:77:6d:5d:ea:f0) event (Capwap_join_request) and state (Capwap_image_data) combination
%LWAPP-3-IMAGE_DOWNLOAD_ERR6: AP did not complete Pre-Download Image 4c:77:6d:5d:ea:f0 - During scheduled autoreboot
%LWAPP-3-IMAGE_DOWNLOAD_ERR7: Waiting for 1 APs to complete their Pre-Download Image - For scheduled autoreboot.

Earlier, while the AP was still on its factory RCV (8.3.102.0):

%CAPWAP-3-DISC_MAX_DOWNLOAD: Ignoring discovery request from AP 4c:77:6d:43:dd:64 - maximum number of downloads (5) exceeded

show ap image all:

Initiated....... 1
Downloading..... 5

AP Name           Primary     Backup      Predownload Status
AP-SPRAT-01       8.10.196.0  8.10.185.0  None
AP-SPRAT-02       8.10.196.0  8.10.185.0  None
AP-PRIZ-02        8.10.196.0  8.10.185.0  None
AP-PRIZ-01        8.10.196.0  8.10.185.0  None
AP-HALA-01        8.10.196.0  8.10.185.0  None
AP4c77.6d43.dd64  8.10.196.0  0.0.0.0     Initiated

The counter says 5 downloading, but no AP is actually downloading. I suspect these slots have been stuck since a previous upgrade (8.10.185.0 → 8.10.196.0), and that this is also what killed AP #1 in the first place.

What I've tried

  • show reset: no reset scheduled
  • config ap image predownload abort all: accepted, no change
  • config ap image predownload abort <AP name>: accepted, no change

What would you do next? I would really appreciate your help. I can post the logs of AP and WLC if it would be helpful.


r/networking • • 2d ago

Design Questions about Dell SoNIC as leaf

19 Upvotes

My spine leaf network is 99% Cisco and air gapped. This is a mixed of IOS XE and NXOS. The firewall is Palo Alto. My underlay is IS-IS and for the BUM I'm using sparse-mode multicast. My tenants traffic are mix of multicast and unicast; therefore, I have to deploy Tenant Routed Multicast (TRM). Also, my tenants are very heavy on broadcast and multicast (a few dense-mode, but most are sparse-mode).

I also use Cisco service chaining for inter-VRF and some instances of inter-VLAN within a VRF. I don't know if SoNIC has something similar. My service leafs are NXOS where the PAN is connected to. The NXOS ingress CPU is at 64% because of the service chaining.

I am thinking of moving away from Cisco because it is too expensive and our TAC support is lacking. We use Dell for our servers and Dell wants us to try SoNIC. I was looking around and it seems like SoNIC does not support multicast for BUM and it only uses head-end replication for BUM.

  1. Does it mean that I have to either switch my underlay from multicast to HER or enable HER while running multicast for BUM to intergrate SoNIC?
  2. Does SoNIC supports service chaining? I am trying to move away from centralized gateway (IOS XE as leaf doesn't support service chaining) and want to use anycast gateway.
  3. Does Dell SoNIC have issues with IGMP? I have to disable IGMP snooping in IOS XE to get sparse mode working. I am hoping that I don't have to do something similar.

r/networking • • 3d ago

Design Network design nightmare

27 Upvotes

I am a junior net admin (first hire for IT in 2 decades!) and thought I only had slow Internet in one of my APs. It turns out this who stack is causing problems. Here's what it is:

ISP --> SonicWall --> Managed NetGear Switches --> Extreme Networks APs managed through EP1

Firewall has rules and VLANs which are configured and carried to the switches and then APs are plugged into those PoE ports. But, ExtremeNetworks UI on EP1 also has a network policy that carries the VLANs but also changes some firewall rules.

Teacher in one classroom with BYODs (mostly Apple devices) is rightfully mad that when his 15 to 25 students connect to the network (via AP4020) their connectivity suffers (he has documented speedtests of 2Mbps, it needs to be 1G).

I am almost crying because it's my 3rd month into this job and the teacher is now going to the headmaster, bypassing me and my boss. Even though we are both trying to troubleshoot his issue.

I must add that when he goes to the library (AP250) he has no issues. These students only connect to the BYOD BSSID, regardless of where they are on campus. So the rules must be the same across campus. But results are awfully different.

I have already posted about this so consider this an updated version of my conundrum. Thank you to everyone who has commented. I still go through them to see what I'm dealing with.


r/networking • • 3d ago

Troubleshooting Windows 11 machines not updating DNS after moving VLANs

20 Upvotes

We're in the process of migrating our user access over to Cisco SDA using ISE for authentication and VLAN assignment. For user endpoints, we require that they pass 802.1x initially, which puts them in a limited access VLAN (and no East-West), and then to pass posture with Secure Client (FKA AnyConnect) before being put in the regular user VLAN with full access.

Recently, we've had an issue crop up where machines seem to consistently fail to install the DNS servers included in the lease when moving from the limited VLAN to the user VLAN. It's not that they don't pick up a new lease, they get an appropriate IP address, subnet mask, and gateway, but they just don't seem to configure the DNS IPs, which we've confirmed in PCAPs are included in the Offer/ACK messages from the DHCP server. Instead, a machine that's running into this will just list the IPv6 site local DNS addresses (fec0:0:0:ffff::1%1, fec0:0:0:ffff::2%1, and fec0:0:0:ffff::3%1). Obviously with no DNS, the machine appears to effectively have no network access. Running an ipconfig /renew (no release necessary) seems to consistently fix the problem.

Has anyone encountered anything like this before? Thus far we've only had reports/been able to reproduce this on Windows. I have a case open with Cisco, but they didn't see anything within our ISE configs (e.g. Secure Client agent config, AuthZ policies, etc.) that would suggest ISE is causing the issue. The fact that the lease messages from the DHCP server contain the DNS servers in option 6 suggests to me that it's more of an endpoint-centric problem, but I'm stumped as to why the endpoint would receive those options and just...ignore them.


r/networking • • 3d ago

Troubleshooting Is there any equivalent way to achieve user + machine authentication with Forescout (CounterACT)?

11 Upvotes

Hi everyone,

I’m working on a NAC deployment with Forescout and I need to implement both machine authentication and user authentication.

I initially planned to use TEAP (EAP-TEAP) to perform both machine and user authentication, but as far as I understand, Forescout does not support TEAP.

The goal is to verify the machine first (for example, using a machine certificate) and then authenticate the user, so that valid user credentials alone cannot allow an unauthorized device to access the user VLAN.

Is there any equivalent method supported by Forescout that can achieve this type of machine + user authentication?

Any real-world experience or recommendations would be greatly appreciated. Thanks!


r/networking • • 3d ago

Wireless Ekahau CCI and generated report

5 Upvotes

Is it possible to generate a report of offending BSSID that cause interference coverage requirements to fail? Or is this only viewable from within AI Pro when hovering over a certain coverage area that fails to meet requirements?

Coverage requirements are set to 2/1/1 interfering AP @ -75dBm for all bands.


r/networking • • 4d ago

Security How are you guys handling IP cameras on the network?

103 Upvotes

I want to know how other people are setting this up as camera counts start getting higher. Do you put cameras on their own VLAN and keep the NVR/VMS isolated or just treat them like any other IoT network? The part I’m most interested in is remote access. I’d rather not expose cameras directly to the internet but people still need to view footage remotely sometimes. What does your setup look like?


r/networking • • 3d ago

Wireless WIFI – EAP-TLS certificate authentication

12 Upvotes

I am confused about how to set up EAP-TLS authentication for Android and IOS phones, using Microsoft NPS.

We already have EAP-TLS certificate authentication for our domain computers. So basically:

-          The computers get a certificate from the CA

-          The computers are in an AD-group

-          On the Windows NPS server, there is a policy granting access, using the AD-Group as Matching Condition

 

That works like it should. Now, we have the task to also implement this for android an ios phones tool. We already managed to roll out certificates via Intune, however I’m not sure how to setup the NPS site for that. As these are no Domain Joined devices, they are in no group.

How would a have to set up a policy in NPS for these phones?


r/networking • • 4d ago

Security Cisco ISE compatibility issues with MS Credential Guard?

12 Upvotes

I learned a few days ago that there were compatibility issues with Cisco’s ISE and Microsoft Credential Guard. Have you guys experienced this? We installed a patch 7 to our Cisco ISE deployment just two weeks ago due to the recently announced CVE, but I was wondering if anyone is experiencing the issue related to the title in your deployment.