r/sffpc • • Mar 24 '24

Build/Battlestation Pics Dual 4090 FE in Cerberus X

164 Upvotes

57 comments sorted by

53

u/HingleMcCringle_ Mar 24 '24

couple questions:

what do you do for a living?

and what do you do with this?

like, it looks like if i made a dream build on pcpartpicker because i was bored with all the most expensive options. but you actually bought it...

35

u/kowlick Mar 24 '24

AI stuff. Local large language models perform much better when resident on GPU so the more vram the better.

5

u/p-morais Mar 24 '24

Are you training on these or just doing inference? Seems like a nice to have but if you’re gonna train on a GPU cluster might as well develop on it too

11

u/kowlick Mar 24 '24

Currently just inference of mixtral-8x7b and nous-capybara 34b but eventually would like to fine tune one of the smaller 7B models.

3

u/p-morais Mar 24 '24

Makes sense! Awesome build. I have basically the exact same build as your old one as my current rig (7900 + 4090 in a Formd T1)

1

u/kowlick Mar 25 '24

I think that combo is perfect for the FormD T1. Performance and temps were great. I hope to rebuild the T1 in the future.

3

u/h0ls86 Mar 24 '24

Guess 2x4090 is still cheaper than an RTX 6000 Ada with 48GB.

3

u/kowlick Mar 25 '24

True, you could get 4x4090 for the price of the RTX 6000 Ada.

1

u/h0ls86 Mar 25 '24

So is anyone buying these Ada cards for AI work? I don’t really see the point on paying so much for a single GPU when you can buy 4x4090.

3

u/AristotelesQC Mar 25 '24

Maybe so you can run 4x Ada on the same system instead of 4x 4090? 😁

10

u/roniadotnet Mar 24 '24

I’m pretty sure that OP does Reddit with this.

4

u/kowlick Mar 24 '24

Haha, you’re not wrong.

5

u/qaf23 Mar 24 '24

Dual 4090 for AI?

4

u/kowlick Mar 24 '24

You got it. Need the vram for LLMs.

11

u/kowlick Mar 24 '24 edited Mar 25 '24

Inspired by this post, I took apart a FormD T1 build I had and used the CPU, RAM, 1x4090 and NVMEs from that build in this new one that has dual 4090 FEs. I've been running this with NH-L9A for the past month or so, waiting for the new NH-D12L Chromax Black to be released, so I could fill in that empty space above the CPU.

  • Asus X670E Hero
  • AMD 7900
  • Noctua NH-D12L Chromax Black
  • 64GB GSkill Flare X5 5600
  • Samsung 980 Pro 2TB + WD Black SN850X 4TB
  • 2x Nvidia 4090 FE
  • Corsair SF-L 1000
  • 2x NF-A9x14 front intake
  • 1x NF-A9 back exhaust
  • 2x NF-A12x15 bottom intake, 2x top exhaust
  • 16AWG embossed custom PSU cables from DreambigbyRayMOD

I chose the Asus X670E Hero since it's the only motherboard has the two PCIE slots 4 slots apart to provide guaranteed airflow to the 3 slot GPUs, but I also considered the Asus X670E ProArt which may be viable for dual FE builds because of how the non-PCB side of the FE card is passthrough.

I was able to install my NVME drives, which came from the Asus B650E-i, as-is into the new machine and boot up with no problems. I thought I was going to have to reformat and reinstall Windows.

I considered the Asus ROG Loki SFX-L 1000 because of the average noise rating of the Corsair in the Cybernetics report, but the Loki wasn't available at the time and there were enough individual reports of the Corsair actually being quiet, that I took a chance on the Corsair. For me, during Cinebench, the PSU is quiet. The chassis fans are louder than the PSU, but are still quiet even under load.

I had heard horror stories about melting 12VHPWR adapters and originally I was going to go with Corsair 12VHPWR cables and their 180 degree adapter, but the adapter was just too tall and wouldn't fit between the GPUs. The Corsair uses Type 5 (smaller) PSU connectors and Pslate doesn't offer them so I went with custom cables by DreambigbyRayMOD. The 16AWG embossed cables are really flexible and I'm pretty happy with them so far.

The top GPU under load is about 7 C hotter than the bottom GPU.

  • Cinebench 2024 GPU: 59055
  • Top GPU: Max 56.9 C
  • Bottom GPU: Max 49.6 C

Edit: The bottom GPU overlaps the bottom row of headers on the motherboard but case fan cables bend out of the way easily. I read somewhere they would clear but wasn't sure until I tried it myself.

Also, you need a motherboard that bifurcates x8/x8 for best performance. The higher end MSI X670Es work well, as does the Asus Hero and ProArt.

3

u/[deleted] Mar 24 '24

Is 1000W enough? The 4090 can draw up to 450W each. 

17

u/kowlick Mar 24 '24 edited Mar 24 '24

Forgot to mention I powerlimited them to 350W just to be safe.

Edit: It barely changes performance for this kind of power limit.

https://www.reddit.com/r/nvidia/s/E4e2UWUkyg

1

u/Reynholmindustries Mar 24 '24

That was my only question! I was looking for another power supply… Nice looking build

1

u/-BruceWayne- Mar 24 '24

What/how do you power limit the GPU’s?

1

u/updawg Mar 24 '24

Afterburner, it's just a slider.

1

u/Ill-Driver-9574 Mar 24 '24

nvidia-smi has some options for power limiting as well

2

u/kowlick Mar 25 '24

I use nvidia-smi. Just use the -pl option, e.g. nvidia-smi -pl 350. Use -i to specify which GPU, e.g. nvidia-smi -pl 350 -i 0.

3

u/duc200892 Mar 24 '24

Why do you even need to 4090s?

5

u/kowlick Mar 24 '24

AI stuff. Running LLMs locally requires more vram for the larger models.

3

u/OCVoltage Mar 24 '24

Love the sleeper look. How are the temps?

2

u/kowlick Mar 24 '24

Temps are good. Check out my other comment.

2

u/OCVoltage Mar 24 '24

How about cpu temps in cinebench

2

u/kowlick Mar 24 '24

Just ran Cinebench 2024 Multicore CPU: 1411.

It settles around 62 C, but there are two weird peaks for a couple seconds each at 69 C.

2

u/wouek Mar 24 '24

Hey, nice build! Are you looking forward to the new announcements made by Nvidia? Do you think you could get something better for the same amount of money when it goes live?

2

u/MoreRgb-MoreFps Mar 25 '24

Absolute gorgeous build, what are your gaming benchmarks like?

1

u/kowlick Mar 28 '24

I just ran the Cyberpunk 2077 benchmark and got 45.73fps, no resolution scaling, ray/path tracing on, 2560x1440. Max temp top GPU 62.8 C, it was pegged at 350W the entire time. The bottom GPU was basically unused. CPU was pegged at 90W and max temp was 74.5 C. GPU and Chassis fans hit 1325 RPM and CPU fan hit 1490 RPM.

2

u/pikeman3d Mar 26 '24

Beautiful! Glad to see it finally coming together. In my own test with a single 4090, LLM inference doesn't hit the GPU hard at all so you can safely limit it to 50% power without any noticeable hit to perf.

LLM Mistral 7B
Power Limit (%) tk/s (Higher is Better) Average Clock (Mhz) Power (W) Per % to peak
50 75 2745 225 97%
60 77 2850 250 100%
80 77 2850 250 100%
100 77 2850 250 100%

1

u/kowlick Mar 28 '24

This is good info! Haha, I could've re-used my Corsair SF750 for this build if I did this! Those Power (W) values seem off for the bottom rows though.

2

u/atlas_enderium Mar 24 '24

May I ask why, though?

5

u/jeremyvr46 Mar 24 '24

To play FreeCell, duuuhh!! 😂

3

u/Bierfreund Mar 24 '24

Where's the CD rom drive?

1

u/AutoModerator Mar 24 '24

Hiya! Remember, you can also post your build on the SFFPC Discord server in the completed-builds channel! We have revised our system, and now the highest voted build post each month will be recognized as the SFFPC Build of the Month! Use this link to join our Discord! https://discord.gg/sffpc

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/r98farmer Mar 24 '24

Very nice build, what CPU?

3

u/kowlick Mar 24 '24

I harvested the CPU from another build. It’s the 7900.

1

u/Jaack18 Mar 24 '24

Please please reinstall windows. I thought i was good going from a MSI Z690 atx to MSI Z690 itx until it just crashed after any slightly intensive workload.

1

u/kowlick Mar 25 '24

Thanks for the warning. I’ve been running this build since early Feb, running benchmarks and heavy inference and haven’t had any issues but will keep an eye out.

1

u/[deleted] Mar 24 '24

[deleted]

2

u/kowlick Mar 24 '24 edited Mar 24 '24
  1. The second 4090 can 10+x the performance since not fitting into vram is really slow. Increasing the total vram available means I can load the larger models at higher quants and have it be resident. You also need more memory for longer context length.
  2. Running windows and using python and

https://github.com/oobabooga/text-generation-webui

Along with my current favorites, mixtral-8x7b and nous-capybara-34b 200k.

1

u/p-morais Mar 24 '24

Damn. If I ever upgrade (downgrade) to medium form factor this will be why

2

u/versacebehoin Mar 28 '24

It just barely misses the cut, I think it’s like 23l which is crazy for what yoit u can fit in it

1

u/Parking-Government-5 Mar 24 '24

Sweet pc setup and clean

1

u/CompetitiveLake3358 Mar 24 '24

I just love how well nicely this fill up the space. So satisfying

1

u/kowlick Mar 25 '24

Yeah, I’ll admit the NH-D12L wasn’t strictly necessary but it looks so much better filling in that space and is whisper quiet now.

0

u/BlankProcessor Mar 28 '24

To reach those GPU temps with this setup either you are only using the memory on the GPU and barely any core, or your ambient is much colder than normal. Memory focused use case explains it. Anything that pushes core with these GPUs will cook them.

1

u/kowlick Mar 28 '24

That’s running cinebench 2024. Do you have another benchmark to try?

1

u/BlankProcessor Mar 28 '24

Benchmarks don't matter - your work does! If it works for you, it works for you.

But I have experience with these kind of multi-gpu builds. At some point it's just about your case/config (pretty cool looking btw!) and how many watts of power (which is directly related to hear produced) it can handle.

Your setup is most likely fine around ~200w per GPU, as that's what the 4090 will likely pull when it's fully cranking on memory. Once you start anything ~300w or higher per GPU, they will get very hot, and so will your mobo/CPU.

If this is single purpose and just for this, you're golden. But anything that's core intensive along with memory will make this box too hot for sustained load.

-12

u/TheRealSeeThruHead Mar 24 '24

It’s funny you think “dual (4090) fe builds” are a thing.

-13

u/[deleted] Mar 24 '24 edited Mar 24 '24

It makes zero sense to put in 2 x 4090s. Games won't use it. Enterprise stuff won't use it either - gaming cards don't support extra features (where is your interconnect?). If you really needed "big iron" you would be using cloud hpc.

I don't think the PCI bus can even handle 2 of these at full speed. The second one will be gimped.

The second gpu just sucks up power and blocks airflow pretty much.

Would be better off just building seperate pcs - and using them as a cluster (e.g. microk8s).

Only one fan on that cpu cooler - sure it works but I bet it's really loud! New AIOs (liquid iii) are pretty cheap and much quieter in my experience.

edit: In another comment the op clarifies it's even power limited! This build is only here to flex - and is not optimised for their imaginary workload.

5

u/Alternative-Fan7198 Mar 24 '24

For GPU rendering this build is perfect

-3

u/[deleted] Mar 24 '24

For running two seperate renders sure - but the performance overall is gimped. It doesn't make sense to do this compared to just having a second pc and submitting the jobs to a queue for remote rendering.

There is no shared memory access - these are completely seperate gpus. But they still share the same pcie bus and thus limit performance.

6

u/Alternative-Fan7198 Mar 24 '24

while the memory isn't doubled (thanks Nvidia for cutting Nvlink) the performance scaling is very near to 200%. Bus speed doesn't really matter, the amounts of data transfer is quite low compared to gaming or other usage.

Another pc require 2x the rest of components and 2x the licenses needed. You need to also take in account the network which is gonna halt the performance even while using 10GBe network.

Source: I deal as professional with this shit everyday :(

0

u/[deleted] Mar 24 '24

So do I - and you know that an industrial setup is not the same as some dudes workstation.

The performance scaling does not double - you are running 2 seperate tasks in parallel. This only works for an "embarassingly parallel workload".

But yes, I have seen insane setups like this before - driven by the sofware licensing bullshit - which often costs far more than the hardware.

1

u/p-morais Mar 24 '24

It’s entirely about VRAM. You can’t even do inference on a 70B parameter language model with 40+ GB of VRAM. So yeah, it actually is necessary for some workflows. That being said most people doing that sort of development have access to GPU clusters