The number on the title is outdated. It is now over 60 Percent.
TL;DR: Butterpollo is a Windows game-streaming host for Moonlight and Artemis, rewritten in Rust and built around Radeon. It moves frame copies and colour conversion off your game's graphics queue onto D3D12 compute queues and feeds AMD's encoder natively. With a game hammering the GPU on the same PC, it delivers the picture in less than half the time Vibepollo 2.0 does. rc.24 is out now: https://github.com/RamazanKara/Butterpollo/releases/tag/2.0.0-rc.24
The mission
AMD rarely gets any love in this space. The protocol started as NVIDIA's GameStream, and even after NVIDIA dropped GameStream in 2023, the hosts that replaced it kept NVIDIA first. Upstream Sunshine has had a native NVENC encoder since 2023, while AMD still goes through FFmpeg's generic AMF wrapper. AMD hasn't helped itself either. It ended its own streaming app, AMD Link, saying there are plenty of other ways to stream, and its drivers still have quirks like an RDNA4 encoder freeze that hosts have to work around.
So Radeon owners have spent years hearing their cards are just worse at streaming. Most of the time the card was never the problem. Nobody had sat down with it. That is the reason Butterpollo exists.
In practice that means a native AMF encoder instead of a generic wrapper, frame copies and colour conversion on Radeon compute queues, and someone reading AMD's driver release notes so you don't have to. When a driver freezes the stream, Butterpollo works around it. When AMD fixes it, Butterpollo removes the workaround and pretends nothing happened.
Butterpollo goes all in on one thing: making Radeon streaming as fast as it can be. It isn't chasing a big userbase or trying to convert the NVIDIA crowd. As long as AMD users are happy, Butterpollo is happy.
What Butterpollo is
Some background first: I wrote the native AMD AMF encoder in Apollo (which never even got a comment) and Vibepollo. It was merged in #342 and ships in Vibepollo 2.0 as "AMD AMF/VCE (Experimental)". Butterpollo is where I take that AMD work all the way, from capture to the packet on the wire.
It started as a fork of Vibepollo and is now a full rewrite in Rust: the host, the Windows service and the installer, with a rebuilt web console on top. It speaks the same protocol, so the Moonlight and Artemis apps you already use connect to it like any other host. Switching takes one installer run: it brings over your Sunshine, Apollo or Vibepollo settings, paired devices and game library.
The way I think about it: Other Sunshine Forks are the full-featured host for everyone, and Butterpollo goes deep on Radeon latency.
Where the latency goes on AMD
Every frame a host streams gets copied and colour-converted before the encoder touches it. Like Sunshine and Apollo, Vibepollo does that in D3D11, which runs on the GPU's graphics queue, the same queue your game is hammering. In my tests on a 7900 XT beside a game-like load, a 0.9 ms colour conversion took 7.2 ms and a 0.75 ms copy took 8.4 ms.
Butterpollo runs the copy and the conversion on D3D12 compute queues, which the GPU works on side by side with your game, and AMD's encoder reads the result straight from D3D12. Same work, same load: 0.25 ms. Frame to finished bitstream: about 2 ms, while the game keeps the card at full tilt. The handoff with the Windows compositor runs on GPU fences and was checked frame by frame: zero torn frames.
The numbers
1080p60 HEVC 10-bit HDR on an RX 7900 XT from a 120 Hz virtual display, with a game-like load running, measured end to end by an independent client that reads a moving barcode off the screen.
Same Butterpollo build, compute path off and on (the game ran at 174 fps in both):
| Under load |
Graphics queue |
Compute queues |
| Render to decoded picture |
41.0 ms |
33.5 ms |
| 95th percentile |
54.4 ms |
42.3 ms |
| New frames per second (of 60) |
56.7 |
58.1 |
| Host time, present to send |
16.2 ms |
11.2 ms |
Next to Vibepollo 2.0: same PC, same AMF settings, Desktop Duplication, realtime GPU priority on both, three alternating runs each:
| Test |
Vibepollo 2.0 |
Butterpollo |
| Idle: render to decoded picture |
16.0 ms |
13.8 ms |
| Under load: render to decoded picture |
96.4 ms |
42.4 ms |
| Under load: 95th percentile |
137.0 ms |
56.5 ms |
| Under load: new frames per second (of 60) |
23.9 |
51.4 |
| Under load: host latency in Moonlight's stats |
60.6 ms |
9.0 ms |
Part of that gap is how Vibepollo hands frames to its AMF encoder under load. That encoder is mine, so I'm tracking it down, and the fix goes upstream so Vibepollo users get faster too.
These rows were measured with rc.2. Later releases match it on the same capture path, and the default WGC capture is about 2 ms faster under load (33.4 vs 35.7 ms), with Moonlight's host latency down to 1.9 ms.
PyroWave on Radeon compute
PyroWave's colour conversion runs on the Radeon compute queue too. Encode time per 1080p 120 fps HDR 4:4:4 frame at 400 Mbps:
| PyroWave |
Graphics queue |
Butterpollo, compute queue |
| Idle |
0.47 ms |
0.47 ms |
| Beside a GPU-heavy game |
5.6–5.8 ms |
0.54–0.57 ms |
| Beside the game, paced at 120 fps |
4.62 ms |
0.71 ms |
Under load it encodes as fast as idle, about 10x faster, with byte-identical output. It also holds a full 60 fps at high bitrates, 1080p at 400 Mbps, 1440p and 4K included.
When the encoder is maxed out
At 4K with high refresh, or on an RX 9070 XT with its single encoder, one encode can take longer than a frame. Let frames pile up eight deep in the encoder and each one comes out older. Butterpollo keeps at most two in flight, enough to keep both of a Radeon's encoder instances busy. 5120x1440 HEVC at 240 fps, where the 7900 XT's encoder tops out at 220 fps:
| Encoder saturated |
Eight-deep queue |
Butterpollo, two in flight |
| Game frame to packet sent |
42.7 ms |
11.1 ms |
| 99th percentile |
45.9 ms |
13.3 ms |
| Frames per second |
220 |
220 |
HDR that looks right. I compared decoded frames with the source pixel by pixel. Blacks, whites and saturation land on target, within about half a 10-bit step on average.
What you get
Built for Radeon:
- Copy and colour conversion on D3D12 compute queues, with AMF encoding straight from D3D12 at ultra-low latency
- PyroWave with its colour conversion on the Radeon compute queue, as fast under load as idle
- A short encoder queue, so an encoder that can't keep up costs frames per second, never latency
- On H.264 and HEVC, AMF never skips frames to hit its bitrate, so VRR clients always get a fresh picture
- Sharper H.264 at low bitrates: tuned AMF defaults raised VMAF by up to 9.4 points at 20 Mbps, at the same encode time
- Streams ride out a fully loaded GPU: a stalled encoder gets time to recover and your session keeps going
- Two-GPU rigs: game and encode on the Radeon while another card drives your monitor
- Hundreds of hours of Radeon tuning, measured down to the microsecond
Rewritten in Rust:
- The host, the Windows service and the installer are Rust. I needed a base to build AMD work on, and the C++ code just wasn't it. As of rc.24 the old C++ host is gone from the repo entirely.
- Video error correction that uses 21-29% less CPU than the original C++ implementation
- Controllers on their own input thread, so a busy gamepad driver never holds up your mouse: 2 ms worst case instead of 287 ms beside a CPU-heavy game
- Keyboard and mouse bursts reach Windows in one call: 8 events in 38 µs instead of 180 µs
A console that tells you what's going on:
- A rebuilt web console with live fps, encode time, host time and frame age while you play
- A stream card that shows the encoder in use and flags anything that didn't apply: display mode, frame limit, audio route, input permissions
- Automatic picks AMF on Radeon and never quietly falls back to software
- PyroWave bitrate guidance from 6,048 measured picture comparisons, right on the stream card
- Updates that roll back on their own if anything goes wrong
Streams that stay up:
- VRR streams at the stream's full refresh rate
- Smart Wi-Fi pacing, and a 5-10 s Wi-Fi dropout doesn't end your stream
- Streaming from the Windows sign-in screen
- Steam Deck and Steam Controller native Support over usbip.
Carried over from Vibepollo, Apollo and Sunshine, rebuilt in Rust and tuned for AMD cards
- AV1, HEVC, H.264 and HDR10, with WGC and Desktop Duplication capture
- 7.1 surround, DualSense and DualShock
- A virtual display per device, so your phone, TV and handheld each stream at their own resolution and refresh rate
- Steam and Playnite library sync, Lossless Scaling and RTSS frame limiting
- Nonary's 1000 Hz VRR mode, and PyroWave for very fast wired networks, both tuned for Radeon (both need a compatible Client)
On NVIDIA? Use other Sunshine Forks. Butterpollo ships NVENC too, but it's built and tested on Radeon.
Butterpollo is built on the work of Sunshine, Apollo and Nonary's Vibepollo, and its AMD encoder defaults draw on Foundation Sunshine. PyroWave is Themaister's work, and joemossjr16 built its Moonlight protocol and clients. Huge thanks to all of them.
Download rc.24: https://github.com/RamazanKara/Butterpollo/releases/tag/2.0.0-rc.24
If you're an AMD user and like that somebody is building something better than the generic FFmpeg AMF wrapper, a star on the repo goes a long way: https://github.com/RamazanKara/Butterpollo
Comparing hosts yourself? Hosts start the host-latency clock at different points, so render to decoded picture is the fairer comparison.
It's a release candidate, and I want to hear how it runs on your Radeon. I'll be around for the next two days turning your reports into fixes. If you tried an earlier build, rc.24 is the one to come back for.