Our CAD renderer can draw 𝟭𝟬𝟬𝗚𝗕 𝗼𝗳 𝗱𝗮𝘁𝗮 on a GPU with only 1GB of VRAM without crashing!
Yes, our system can handle 10-100GB of draw data through only 1GB of VRAM. But if we have to access all of that in a single frame, we constantly thrash the GPU cache and move data from the CPU to the GPU (because no data can stay consistent across frames).
People tend to overestimate PCIe upload speeds. Normal PCIe speed for large contiguous transfers sits around 10-30GB/s depending on the hardware. That means pushing 10GB takes 𝟬.𝟯 𝘁𝗼 𝟭 𝘀𝗲𝗰𝗼𝗻𝗱 𝘁𝗼 𝗳𝘂𝗹𝗹𝘆 𝘁𝗿𝗮𝗻𝘀𝗳𝗲𝗿. In a CAD application, that completely destroys the interactivity engineers rely on.
𝗪𝗵𝗮𝘁 𝗰𝗮𝗻 𝘆𝗼𝘂 𝗱𝗼 𝗮𝗯𝗼𝘂𝘁 𝗶𝘁?
If your app can spare the time to preprocess data, The standard solution is spatial acceleration structures paired with coarse-grained cpu culling to stream only relevant blocks into VRAM. And to handle extreme zoom-outs where the entire model is visible at once, you pair this with a Level of Detail (LOD) system.
𝗕𝘂𝘁 𝘄𝗵𝗮𝘁 𝗶𝗳 𝗮 𝘀𝗶𝗻𝗴𝗹𝗲 𝗳𝗿𝗮𝗺𝗲 𝘀𝘁𝗶𝗹𝗹 𝗲𝘅𝗰𝗲𝗲𝗱𝘀 𝘆𝗼𝘂𝗿 𝗩𝗥𝗔𝗠 𝗮𝗹𝗹𝗼𝗰𝗮𝘁𝗶𝗼𝗻?
We recently found ourselves in a situation where we had to push a 1.3GB non-preprocessed frame through a strict 1GB (minimum requirement) VRAM allocation.
𝗧𝗵𝗲 𝗴𝗼𝗼𝗱 𝗻𝗲𝘄𝘀? Our auto-submit mechanism handled the overflow beautifully, automatically splitting the draw into two submits, making sure all relevant data is versioned and cached in VRAM
𝗧𝗵𝗲 𝗯𝗮𝗱 𝗻𝗲𝘄𝘀? it still meant we were pushing data across the bus every single frame: The PCIe transfer took ~100ms, while actually rendering the data took only 8ms!
To fix this bottleneck, we have started aggressively packing our data using bitfields and introducing new instancing mechanisms into our renderer.
If my calculations are correct, by the end of next week we’ll be down to ~600MB of data representing the exact same view. This allows the dataset to reside persistently in VRAM, avoiding constant copies and keeping the viewport interactive.
Final Note:
Remember that the foundation of any out-of-core rendering system relies on CPU being aware of what the GPU has in it's VRAM cache and making decisions about what to evict/free, and most importantly ensuring correct synchronization to avoid unwanted GPU crashes (will post about this in more detail soon, so stay tuned)
This writing mostly applies to non-UMA devices with discrete GPUs.