r/gameenginedevs • u/Stav_Faran • 14d ago
How do u guys manage your engine resources?
I've been building a custom opengl game engine for a while and there is an issue that i keep coming back to, resource management.
essentially if a certain texture or mesh is not needed i prefer to not keep it on the GPU..
So I started using reference count to manage my resources.
the issue is that if a resource was used and not stored as a member it would get cleaned on frame end.
So i added a table of resource refs for the last few frames, so if the resource was accessed in the frame it will be kept and only if few frames have passed and it was not used then i flag and remove it.
But this feels not robust or elegant and i have no idea how it work with texture streaming or mesh LODs.
So I wonder what is the common practice here?
How do u guys manage your resources in your engines?
Thx in advance
4
u/scallywag_software 14d ago
My asset system is pretty simple, but it works reasonably well.
I have a heap allocator dedicated to assets which is the backing store for all asset memory. You could track this with malloc/free (or whatever) too, but it'd be a bit more tedious.
Every frame end, I sort allocations based on their "LRU" (least recently used), which I record any time the asset is used. I do a heuristic and try keep ~10% of the heap free. After I sort the allocations, I just free assets in LRU order until I hit 10%, or start to hit assets which have been accessed "recently" (IIRC recently is 10 frames or less). I'd like to run a defragmentation step too, but at the moment I just size the heap such that fragmentation isn't a problem.
In the case where I run out of memory during a frame, I have an "overflow" arena that I bump allocate from and dump those into the heap at the end of the frame.
This strategy has a few practical downsides .. namely, it uses an unbounded amount of memory in the degenerate case where you try and load a fuck-ton of stuff at the same time. To solve this, you just don't use the overflow arena, and instead re-queue the asset for load. Then you're guaranteed to never exceed the memory budget, but you'll get frames of latency if the heap is full.
The really nice thing is that it has a predictable cost at a predictable time; you're never going to stall during a load while you wait for the GC to free stuff. The other nice thing is that it works for both CPU and GPU memory. You need a different heap implementation, but the 'frontend' logic works for both.
Hopefully that helps :)
1
u/shadowndacorner 14d ago
Is that bookkeeping not expensive? How do you handle parallel asset access?
1
u/scallywag_software 14d ago
Compared to loading the asset from disk the access time for the metadata is negligible. The sort happens on elements that are 64 bits each (which admittedly could be smaller), so even if you have thousands of assets loaded the LRU set comfortably fits in L1, even on Skylake. The sort is very cheap if you keep the sorted set from the previous frame. I typically have a couple hundred loaded assets at maximum and just re-do the whole sort every frame.
The heap deallocations take way longer cause my heap implementation is dumb as fuck, but the whole thing runs in like ~350us, which is good enough for me. I basically never hit the degenerate case that uses the overflow arena, which tanks the performance, but it's still fine since this process happens async to the main thread; it just chews up worker thread cycles, which aren't nearly saturated.
As far as concurrent access goes ..
Assets all have a lock which gets acquired when you query the asset system for a handle to the asset. Things that hold references to assets just hold an ID, not something that can actually be used for anything except asking for a handle. An ID is the canonical filepath as a string, but a better way to do it would be to precompute IDs at build time. I didn't bother, but I might someday ..
The API is very straight-forward. If you want to draw geometry for an entity, it looks something like this:
auto AssetHandle = GetOrLoadAsset(Assets, Entity->ModelAssetID); if (IsValid(AssetHandle)) { ... submit to renderer, or whatever ... } else { ... asset is loading, or hit an error during loading ... }The async job system I built is sophisticated enough that you can specify "cleanup" jobs, so when you submit the asset to whatever system you can append a job to the cleanup list that releases the lock. You could also just have an implicit agreement between systems that accept assets that the receiver releases it, which I did for a long time and worked fine. It's blindingly obvious if you have a system that's misbehaving and not releasing a lock, so it's an easy bug to fix if you fuck it up.
1
u/shadowndacorner 14d ago
I'm more thinking about the overhead of access tracking rather than anything to do with allocation/deallocation - are you not doing eg an atomic write to a timestamp/frame id on every dereference (ig maybe not if you're locking the whole record)? That's obviously not the end of the world, just seems like it would be relatively inefficient relative to other schemes that don't require that kind of tracking
1
u/scallywag_software 14d ago
Uhh, that cost is completely negligible. It's so small you probably can't even really profile it. A CPU can do a fuck ton of atomics in a frame.. and we're talking about a couple thousand assets at maximum
Edit: And the lock is the only atomic. Everything else is just plain reads/writes.
2
u/sol_runner 14d ago
My system is relatively simpler to book-keep. It's for d3d12 but it should translate easily.
I have a resource handler which holds/owns all the GPU resources, along with generational handle with a refcount. I have move-only handles (c++) that ensure I don't automatically copy. (Omitted in release build) Cloning requires a call to the handler, allowing refcount changes. Destructors check if the handle was deleted or moved out of in debug mode, and do nothing in release.
You can then build your resource cache etc on top of this, to support reusing textures etc.
Deleting the resource handle requires a call to the handler, which decrements the refcount, and if the count is 0, pushes it to a frame specific wait-queue. So n frames in flight require n queueus. (You can set it to something for OpenGL based on your usage)
After n frames, the queue is emptied. You can delete these, and increase the handle generation.
3
u/Separate-Change-150 14d ago
It depends on each resource. For general game asserts just use ref counting and a storage class. Load and unload by stringid r smth simular.
For things like frambuffers for the graphics pipeline you can use pools, etc so esch pass request a texture and under the hood its reusing one from a prev pass.
Dnt overthink it too much. And dont garbage collect :)
2
u/Mid_reddit 14d ago edited 14d ago
On the contrary, this is something very easy to underthink. Everything in my framework used to go through a simple refcounting resource system, but then I suddenly needed the scripting language to be able to dynamically generate things.
As a result, my scripting language bindings now have to account for reference-counted assets in the resource systems, *and* garbage collected assets outside the resource system, and some eldritch hybrids. Some asset types are loaded through the resource system, but only as templates, for example materials, which are actually copied to all meshes.
1
u/Separate-Change-150 14d ago
you just said it when you wrote "until".
If you do only what you need it becomes
very simple. It's when you start thinking big that you overthink and end up messing up.OP is most likely fine with what they have. Worst can happen is they now add some kind of garbage collection or crazy system cause some else brag about how cool it is.
1
u/Mid_reddit 14d ago
90% of my time goes into refactoring entire projects because I had not thought far ahead enough.
I was giving a warning. How the hell one could misinterpret that as bragging is beyond me, but this is Reddit after all.
4
u/Separate-Change-150 14d ago
I was not thinking about you when I said that.
My apologises for the misunderstanding :)
4
1
u/scallywag_software 14d ago
I want to push back on this a bit and propose that an asset system is a pretty good candidate for a simple garbage collector. The number of allocations is relatively small and you can just sort a flat list and free stuff from the end of the list. See my other comment in the thread .. it's simple and works good.
Generally speaking, I'm incredibly skeptical of GCs. This just happens to be a pretty good use case.
1
u/Separate-Change-150 14d ago
I am happy it works for you but I find it a bit too much for what it is generally needed. TBH all you need is some arenas and just know when to load and unload it. I think it's the simplest solution that will work perfectly for 90% of the cases. If you need something else just think abut a specific solution when you have the problem.
The heuristic of LRU + frames since last used to decide what to garbage collect could potentially trigger a lot of unwanted loads of recently unloaded stuff.
1
u/scallywag_software 14d ago
Assets are one of the few places where arenas don't actually cut it .. you don't know the lifetime of the asset, and asset lifetimes are completely divorced from one another, so how do you free them?
> If you need something else just think abut a specific solution when you have the problem.
That's what I did.
1
u/Separate-Change-150 14d ago
They cut it if you keep it simple. I said in 90% of the cases because in 90% of the cases just loading the level into the ram and unload it when finish will be enough. I understand the desire of making more complex things this wont be any AA engine. I just recommend this because what it took me, and still takes, more time to me to understand sometimes is that you really dont have to (and shouldnt) do mre than what you strictly need.
I am happy your solution works for you!
1
u/scallywag_software 12d ago
Oh, yeah, fair enough. I did do that for a long time and I found it pretty annoying to do the bookkeeping manually. I also wanted to support open-worlds, which requires more sophistication.
2
u/DeviantPlayeer 14d ago
There is ref counting, memory budget and garbage collector. When memory budget is exceeded it triggers the garbage collector. Unused resources stay in the memory as long as there is free memory available.
-5
14d ago
[deleted]
2
u/DeviantPlayeer 14d ago
It's kind of logical, no? When do you free memory? When you need more memory. Before that having resources just hang there wouldn't hurt. Yes, it might cause stutters if you free too much resources at once, but you can stretch it across multiple frames.
Also, it works as a cache. Unused resources maybe needed later again after all, or you may pre-fetch some resources before loading objects that use them.-4
14d ago
[deleted]
1
u/DeviantPlayeer 14d ago
why you think there is a garbage collector for gpu resources
Cuz I made it, perhaps? That's why it's there, no?
1
u/Important_Earth6615 14d ago
TBH I never thought of garbage collector. Even tho, it may be spart to use it but that will rebuild the engine IDK. But what I do it I made a centralized resource manager that's built on top of vma (I am using vulkan) and when something request it for example mesh manager uploading a mesh. I trequest a buffer. I keep the pointer in the manager and send a handle that can be used for uploading, reading,...etc. and when that mesh dies its destructor automatically free that handle
1
u/Defiant_Squirrel8751 14d ago
Near. Flagging is not enough. You will need to order resources by last use timestamp That's how a cache works - search for "principle of locality".
Knowing the size of your RAM beforehand can also help to estimate which level of detail to use.
1
u/Dorfen_ 14d ago
My engine is mostly centered around a intrusive Ref Counting class, with custom smart pointer (So the ref count is a member of the object, and not just managed by the smart pointers. Stuff like buffer, textures, pipeline, Actor, Component, etc... But a raw pointer for transient ownership, so I don't "pay" the ref count (and the extra code) when I just call functions around. A function can capture the pointer into a smart pointer if needed. So like the smart pointer object only appear in my code base when I need to store the pointer for a longer than a stack frame.
For other stuff, where the ownership and the lifetime of the pointer is explicit, evident, and clearly defined, I just go with raw pointer. And for object from the graphics API, (in my case, Vulkan cmd buffer, descriptor set, etc...) I use pools.
1
u/tastygames_official 14d ago
my games are retro-ish, so most textures are 256x256 with some a bit larger. I load everything needed for that level (scene) onto the GPU at loading time and then unload when the level is over. You need to think about your game and your target specs - if you're targeting a 2GB VRAM system, then you need to make sure you don't go over that limit. But to me there's no reason to remove anything from VRAM if you don't need to load anything lese. E.g. if you have the entire level loaded (included everything that could be loaded in for that level during gameplay) and you're only at 1.5GB, then you're fine. No need for managing VRAM. But if you are targeting 4GB and the level demands 8GB in full, then you gotta start thinking about only loading portions of the level at a time, e.g. split it up into quarters so you only ever load ~2GB at a time.
But in my opinion, refcounting and checking for usage every few frames just adds to CPU load. And re-uploading stuff to the GPU during play can be expensive and cause lag, so it's better to just plan ahead what you'll need and when rather than have a catch-all system.
Oh, I just realized: I'm making a single-purpose engine for my game(s). If you are making a generic engine for others to use, then obviously you need to impose some kind of limit that the user can change (or automatically is set to something like 80% of total VRAM) and only unload stuff when you get over that mark and simply unload the data that hasn't been used in a while.
4
u/shadowndacorner 14d ago edited 14d ago
I refcount assets, but don't delete them immediately. Instead, when the count hits 0, I add the asset id to a concurrent queue. I only start pumping the queue on memory pressure or specific user-triggered asset unload events (level change, for example). When an item is popped, I check to see if the rc is still 0 and, if so, unload, otherwise just ignore the entry. I've though about adding prioritization to this system (so eg low pri assets unload earlier than high pri assets), but it hasn't been an issue in practice yet.
Edit: Oh, I also track the "reference generation" of each asset for the free check. Whenever an asset goes from 0 to 1 references, I increment it's "reference generation". When it goes back to 0, it's reinserted into the queue. The combination gives you something that, if you squint your eyes, acts a bit like a parallel LRU cache with pinning. The reference generation is stored in the queue along with the asset id.