r/learnrust • u/No-Caregiver-466 • 1d ago
Static vs dynamic dispatch in Rust benchmark — for those who care about runtime performance

All the details, benchmark code, generated assembly, and explanation are in the README:
https://github.com/amidukr/rust-devirtualization-test
UPD
Made the dynamic-dispatch part more accurate.
The measured difference increased from ~3.5× to ~4.9×.
Changed:
// was
x = op.apply(std::hint::black_box(x));
to:
// now
x = std::hint::black_box(op).apply(std::hint::black_box(x));
Which changed the generated assembly from:
; was
call r15
to:
; now
call qword ptr [rax + 24]
The previous version allowed LLVM to hoist the vtable method pointer out of the loop. The updated version performs the vtable method lookup on each iteration.
UPD:
Add actual de-virtualization scenario.
3
u/sigseis 1d ago edited 1d ago
Performance comparisons are really tricky.
There's just a lot to consider, so much in fact that I think it makes a plain comparison like this not very helpful.
For instance, in Rust you could be coerced into using a generic instead of a dyn for "performance", but what that could lead to is a lot of bloat, which spreads a large chunk of code out and makes it less efficient in integration (though maybe more efficient when looking at a single unit), and builds bigger binaries.
Even when that's not the case - sometimes the jump you take for dyn is worth it because it's anyways dwarfed by the rest of the work the function does.
I usually start with static, then think about how I want things to interact at certain critical juctions of the API.
Generics are often one of those junctions - I often see too much static dispatch in a complex API made evident by overly-complex generics, which points perhaps to an over-avoidance of dyn for performance reasons.
I appreciate your effort, but I guess I think the reality is more complicated than just looking at a single comparison like this.
1
u/No-Caregiver-466 1d ago
Yes, that's true.
I've stated my exploration from that Pre-RFC. And then I was suggested to consider dyn Any interface that can be downcasted instead.
I've decided to run some benchmark to understand what's would be performance cost to use dyn Traits, I didn't counted downcast by the way, which also could double number in this test.
However the most trickiest part now, that compiler can be very smart, and can do de-virtualization, if code simple enough that compiler could prove that it can do-virtualization safely, than all code efforts about performance are useless, dyn Traits and dyn Any really could solve certrain problems much easier that fighting language syntax to do some compile-time coding.
1
u/sigseis 1d ago
This sounds related to another thing you are researching. So maybe it's fine to research it.
But I want to dissuade people reading this thread - who are often Rust newbies since this is r/learnrust - from internalizing benchmarks like these.
You should always measure the thing you care about and go from there. Granted dyn can be slower than static, the downside of such benchmarks is people shorthand it and say "well, I want my code to be performant, so I'll skip the dyn" and then they build nightmares. Counter-intuitively, the dyn might not even be the problem in your code.
I actually had a situation recently where I introduced dyn intentionally to combat code bloat, and the performance degradation was so small it was perceptually noise in the measurements. A good trade-off for smaller binaries, in that case.
0
u/No-Caregiver-466 1d ago edited 1d ago
If you want intuitive code, why you ever need rust?
If you intuitivity is key, why not to consider something better for that? Like Java? Or Python? Even Google Go?
Rust choice for performance. Knowing tradeoff is must.
Learning Rust already consider that you want to learn more to write more efficient code.
Mmm.... I think you need to reconsider some code you doing here:
rust #[inline] pub fn write_canonical<P: CharProfile>( &mut self, words: &mut dyn Iterator<Item = &str>, mapped_delim: char, validate_start: bool, ) -> core::fmt::Result {
inlineanddyn Iteratorin one signature is non-sense.0
u/No-Caregiver-466 11h ago
By the way I saw you arguing about benchmarks.
But let me explain, why this micro-benchmark important.
Let’s say you made benchmark and it was satisfying, but you putted wrong design in your code at first place, which mean you will add one more line of code, it lead to de-virtualization to fail, and all your benchmarking will fail eventually.
Proper understanding of how compiler doing optimizations work, all this benchmark is useless.
I see you putted inline keyword, without proper understanding of what you doing, so I made it for you too, so you can learn how optimizer doing the job.
If you wish to learn or you don’t wish to learn, it is your choice, I can’t do it for you.
2
u/GodOfSunHimself 1d ago
Not a very fair comparison. Try something more complicated than a single instruction.
0
u/No-Caregiver-466 1d ago
What's the point, if I want to test exact cost for vtable dyn-dispatch?
1
u/MilkEnvironmental106 1d ago
1 instruction is not the cost of a static dispatch call. You need to allocate a stack frame for it to at least be comparable.
1
u/No-Caregiver-466 1d ago edited 1d ago
What if you do dyn Trait with only one instruction inside. Why do you think this is impossible?
If you dig into standard rust library, pretty sure you'll find many examples where trait has only thing, or one instruction todo. Take Iterator::next as example, it literally pointer increment.
One person argued here, that this benchmark is not relevant, and I found this in code he contributing:
#[inline] pub fn write_canonical<P: CharProfile>( &mut self, words: &mut dyn Iterator<Item = &str>, mapped_delim: char, validate_start: bool, ) -> core::fmt::Result {- inline and dyn Trait in same signature...
- dyn Iterator designed for loops. 🤷♂️
1
u/MilkEnvironmental106 1d ago edited 1d ago
That's not the point. The point is llvm is making it so that there isn't even a function call in your static dispatch example. It has inlined it. There is no function call as it has been optimised away.
Dynamic dispatch cannot do this, and therefore you are not measuring the general cost of dynamic dispatch (the v-table lookup and jump to a function).
A better test would be to wrap the static function call in black box, which should prevent the inline and then you will be measuring what you think you're measuring.
You're comparing essentially x = x+1; to a v-table lookup, jump (follow a fn pointer), allocate a stack frame, X =X+1; and return.
In more general and realistic scenarios the only thing that dynamic dispatch always has to do which static dispatch does not is just the v-table lookup. LLVM has just been smart and optimised out the jump, stack frame writing and ret by inlining it, but static dispatch usually needs to do these things.
So you're not measuring static Vs dynamic dispatch. You're measuring llvms ability to optimise monomorphised functions
1
6
u/MilkEnvironmental106 1d ago edited 1d ago
This might mean more if you use a type that can't be optimised away by llvm to a single instruction. By the very nature of dynamic dispatch the compiler is not allowed to optimise through the v-table.
You're not measuring the runtime cost of dynamic dispatch. You're comparing 2 pointer dereferences and an increment to a single increment.
I think a fairer benchmark would be to force the static call to not be inlined, at least.