Devsh Graphics Programming’s cover photo
Devsh Graphics Programming

Devsh Graphics Programming

Gdańsk, Województwo Pomorskie 349 followers

Whole teams of expert GPU and Rendering engineers, today and without HR lead time

About us

Devsh Graphics Programming is a company based out of Ul. Lipuska 36, Gdańsk, Województwo Pomorskie, Poland.

Website
https://www.devsh.eu
Headquarters
Gdańsk, Województwo Pomorskie

Employees at Devsh Graphics Programming

View 9 employees at Devsh Graphics Programming

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

Updates

  • Devsh Graphics Programming reposted this

    500+ graphics engineers from around the world will be at the GPC 2026. These companies are making it happen: Nvidia Intel AMD JangaFX Keen Games Devsh Graphics Programming @Superluminalbv Qualcomm Arm Nilo SEED (Electronic Arts) If you've not heard about GPC yet, you'll get a good mix of industry veterans and a very high level of talent that shows up every year. Last year, we not only sold out the event, but also every single hotel room in the city! Hope to see you there. November 17–19. Breda, Netherlands.

    • No alternative text description for this image
  • Instead of having your $250'000 engineer spend $250'000 on tokens, spend $2500 on our Corporate Training offering instead.

    Want to learn 𝗖++ 𝗳𝗼𝗿𝘄𝗮𝗿𝗱 𝗽𝗿𝗼𝗴𝗿𝗲𝘀𝘀 𝗴𝘂𝗮𝗿𝗮𝗻𝘁𝗲𝗲𝘀? tempted to ask AI for an easy explanation? we tried "Gemini Pro" and "Claude Fable", not only they weren't able to teach it properly. their explanations were completely wrong and hallucinated. I'm ashamed to even post them here. Instead of going on circles promoting. watch this 1:30h masterclass by an expert "human" Olivier Giroux: https://lnkd.in/d_tpuRHv Please Please, before asking AI for anything. spent at least a minimum of 5 minutes thinking about it yourself. There is NO learning without failure and struggle! Go in the c++ spec. fail to understand the syntax, ask your mentors point of view. search for other humans perspective and writings.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • If only #Nabla had a Linux build then we could probably run this benchmark and show Mesa being superior to many Windows drivers. I've seen the code, its correct, the proprietary drivers... who knows? https://lnkd.in/e9gwZWQQ

    What's the best way to 𝘀𝘁𝗿𝗲𝗮𝗺 𝘆𝗼𝘂𝗿 𝗶𝗺𝗮𝗴𝗲𝘀 𝘁𝗼 𝘁𝗵𝗲 𝗚𝗣𝗨 using modern graphics APIs? We benchmarked streaming thousands of 128x128 tiles for our Virtual Texturing system with different methods:  • Regular copyBufferToImage   • Using compute shaders to do the copy  • Using the new host_image_copy extension in Vulkan (promoted in 1.4 version)  • Using staging memory in different memory spaces (VRAM ReBAR, VRAM Staging, RAM, UNCACHED RAM) 𝗖𝗼𝗻𝗰𝗹𝘂𝘀𝗶𝗼𝗻? • Even though vulkan specification mentions host_image_copy could be potentially faster than staging method, we saw 2-3x slowdown on most of the hardware we tested on. Even with the MEMCPY_BIT  • Using a compute shader vs copy command of graphics queue didn't seem to make any meaningful difference. (Only use compute if you want to decode + copy in one go)  • 𝗙𝗶𝗻𝗮𝗹 𝗩𝗲𝗿𝗱𝗶𝗰𝘁: Just use an staging buffer, and the best place for your staging area seems to be in device accessible RAM. (make sure it's write-combined memory) Thanks to Abbas Garousi for writing the benchmark.

  • The 2D CAD view of your survey app can't hog all the VRAM, because ML and those sweet sweet 3D point clouds need it, so how do you do Out of Core rendering? https://lnkd.in/dMfcYkjh

    Our CAD renderer can draw 𝟭𝟬𝟬𝗚𝗕 𝗼𝗳 𝗱𝗮𝘁𝗮 on a GPU with only 1GB of VRAM without crashing! Yes, our system can handle 10-100GB of draw data through only 1GB of VRAM. But if we have to access all of that in a single frame, we constantly thrash the GPU cache and move data from the CPU to the GPU (because no data can stay consistent across frames). People tend to overestimate PCIe upload speeds. Normal PCIe speed for large contiguous transfers sits around 10-30GB/s depending on the hardware. That means pushing 10GB takes 𝟬.𝟯 𝘁𝗼 𝟭 𝘀𝗲𝗰𝗼𝗻𝗱 𝘁𝗼 𝗳𝘂𝗹𝗹𝘆 𝘁𝗿𝗮𝗻𝘀𝗳𝗲𝗿. In a CAD application, that completely destroys the interactivity engineers rely on. 𝗪𝗵𝗮𝘁 𝗰𝗮𝗻 𝘆𝗼𝘂 𝗱𝗼 𝗮𝗯𝗼𝘂𝘁 𝗶𝘁? If your app can spare the time to preprocess data, The standard solution is spatial acceleration structures paired with coarse-grained cpu culling to stream only relevant blocks into VRAM. And to handle extreme zoom-outs where the entire model is visible at once, you pair this with a Level of Detail (LOD) system. 𝗕𝘂𝘁 𝘄𝗵𝗮𝘁 𝗶𝗳 𝗮 𝘀𝗶𝗻𝗴𝗹𝗲 𝗳𝗿𝗮𝗺𝗲 𝘀𝘁𝗶𝗹𝗹 𝗲𝘅𝗰𝗲𝗲𝗱𝘀 𝘆𝗼𝘂𝗿 𝗩𝗥𝗔𝗠 𝗮𝗹𝗹𝗼𝗰𝗮𝘁𝗶𝗼𝗻? We recently found ourselves in a situation where we had to push a 1.3GB non-preprocessed frame through a strict 1GB (minimum requirement) VRAM allocation.  𝗧𝗵𝗲 𝗴𝗼𝗼𝗱 𝗻𝗲𝘄𝘀? Our auto-submit mechanism handled the overflow beautifully, automatically splitting the draw into two submits, making sure all relevant data is versioned and cached in VRAM 𝗧𝗵𝗲 𝗯𝗮𝗱 𝗻𝗲𝘄𝘀? it still meant we were pushing data across the bus every single frame: The PCIe transfer took ~100ms, while actually rendering the data took only 8ms! To fix this bottleneck, we have started aggressively packing our data using bitfields and introducing new instancing mechanisms into our renderer. If my calculations are correct, by the end of next week we’ll be down to ~600MB of data representing the exact same view. This allows the dataset to reside persistently in VRAM, avoiding constant copies and keeping the viewport interactive. Final Note: Remember that the foundation of any out-of-core rendering system relies on CPU being aware of what the GPU has in it's VRAM cache and making decisions about what to evict/free, and most importantly ensuring correct synchronization to avoid unwanted GPU crashes (will post about this in more detail soon, so stay tuned) This writing mostly applies to non-UMA devices with discrete GPUs.

  • We are proud to announce our official sponsorship of the Eurographics Symposium on Geometry Processing (SGP) 2026, hosted this year at the University of Bern Excited to be among the leading experts, researchers, and students in Geometry Processing. Are you attending SGP this year? Let's connect! Our own Krzysztof Szenk is going to be there. 👕✨ Make sure to grab your exclusive glow-in-the-dark nabla t-shirt! Let us know if you will be there in the comments below! 👇 #SGP2026 #GeometryProcessing

    • No alternative text description for this image
    • No alternative text description for this image
  • Devsh Graphics Programming reposted this

    What happens when a tiny rounding error on your GPU shifts a massive 𝗖𝗔𝗗 𝗺𝗼𝗱𝗲𝗹 𝗯𝘆 𝗮 𝗳𝘂𝗹𝗹 𝗺𝗲𝘁𝗲𝗿? If you work with massive, georeferenced CAD datasets, you know that scale and precision are non-negotiable. Naive usage of 32-bit floats (single-precision) to represent your 𝗪𝗚𝗦𝟴𝟰 or 𝗘𝗣𝗦𝗚 𝟯𝟴𝟱𝟳 coordinates could 𝗹𝗲𝗮𝗱 𝘁𝗼 𝟬.𝟱 𝘁𝗼 𝟮 𝗺𝗲𝘁𝗲𝗿𝘀 𝗼𝗳 𝗲𝗿𝗿𝗼𝗿! Switch to 64-bit and you get error on the scale of 𝗻𝗮𝗻𝗼𝗺𝗲𝘁𝗲𝗿𝘀, that's more than enough for applications representing points on Earth's surface. But that's easier said than done for a 𝗚𝗣𝗨-𝗗𝗿𝗶𝘃𝗲𝗻 𝗥𝗲𝗻𝗱𝗲𝗿𝗲𝗿, where some GPU vendors don't report float64 support [*cough* Intel Arc *cough*] 𝗪𝗵𝗮𝘁 𝗱𝗶𝗱 𝘄𝗲 𝗱𝗼? We didn't compromise, we emulated IEEE754 ourselves in HLSL using int64 (Thanks to Przemyslaw Pachytel).  𝗜𝘁'𝘀 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲! Take a look: https://lnkd.in/dwsdHp7F Even if you find a way to work with fp32 and make it "good enough" for static representation, you still need to worry about operations where errors propagate rapidly: like root finders, polynomial solvers, differential equation solvers, and curve fitting. That's why we always do some form of error analysis on our important functions that deal with geometry, whether we use fp32 or fp64 (float or double). (Note: We do still rely on standard fp32 arithmetic for our complex SDF functions in screen space. The Gallery below shows some of the fun bugs we had to deal with) Recommended Resource for more technical readers: Read "What Every Computer Scientist Should Know About Floating-Point Arithmetic" by David Goldberg P.S. We specialize in building high-performance, GPU-driven renderers for fields where precision matters. See what else we are building at DevSH: https://www.devsh.eu/ #GPUProgramming #ComputerGraphics #GIS #CAD #Rendering

  • Devsh Graphics Programming reposted this

    𝗧𝗿𝗮𝗻𝘀𝗽𝗮𝗿𝗲𝗻𝗰𝘆 in 2D Rendering isn't as straightforward as you might think. Drawing an object with the 𝗽𝗿𝗲𝘃𝗶𝗼𝘂𝘀 𝗼𝗯𝗷𝗲𝗰𝘁'𝘀 𝗰𝗼𝗹𝗼𝗿? But why!? Building 2D GPU-Driven renderers is a unique challenge. Instead of standard meshes, you are directly rendering lines, curves, UI components, and vector graphics. If you want transparency on these 2D components while leveraging the GPU's blending unit, you have to find a way to prevent it from blending with itself. 𝗧𝗵𝗲 𝗵𝗮𝗿𝗱𝘄𝗮𝗿𝗲 𝗹𝗶𝗺𝗶𝘁𝗮𝘁𝗶𝗼𝗻: The GPU doesn't inherently know which fragments belong to the same UI Component or Polyline. The graphics pipeline will just blindly blend overlapping pixels. To solve this, we built a custom software blending algorithm (battle tested in #n4ce's 2D survey renderer) Here is how it works:  • We use an object-aware offscreen texture to accumulate alphas, rather than outputting to render target directly.  • When do we finally output to the target? Only when we are absolutely certain we are DONE drawing the current object.   • If the next object gets drawn on top, we render it using the previous object's color and resolve!   • A fullscreen pass at the end resolves the colors of anything that was accumulated but not yet written to the render target. The logic is sound, but hardware execution introduces a caveat. While a GPU guarantees primitives will blend in the order they were submitted, 𝗶𝘁 𝗱𝗼𝗲𝘀 𝗻𝗼𝘁 𝗴𝘂𝗮𝗿𝗮𝗻𝘁𝗲𝗲 your fragment shaders will process them in that same order. To prevent race conditions, this algorithm relies on 𝗙𝗿𝗮𝗴𝗺𝗲𝗻𝘁 𝗦𝗵𝗮𝗱𝗲𝗿 𝗜𝗻𝘁𝗲𝗿𝗹𝗼𝗰𝗸. This hardware feature enforces strict synchronization, ensuring reads and writes to our offscreen alpha accumulation texture happen in the exact order the draws were submitted. We’ve presented this algorithm at Vulkanised 2024. (Link to presentation in the comments) #GPU #Rendering #GraphicsProgramming #Vulkan

  • This is what's already been done, just imagine what we're currently cooking. https://lnkd.in/eyC9uW-A

    𝗜𝗻 𝗖𝗔𝗗 𝗿𝗲𝗻𝗱𝗲𝗿𝗶𝗻𝗴, 𝗳𝗶𝗹𝗲 𝘀𝗶𝘇𝗲 𝗱𝗼𝗲𝘀𝗻'𝘁 𝘁𝗲𝗹𝗹 𝘁𝗵𝗲 𝘄𝗵𝗼𝗹𝗲 𝘀𝘁𝗼𝗿𝘆. A 1GB land survey DTM model on your disk doesn't sound scary. But when you decompress it for drawing, it can easily balloon into 𝟭𝟬𝗚𝗕 of system RAM. How do you render that when the user’s GPU only has 𝟮𝗚𝗕 𝗼𝗳 𝗩𝗥𝗔𝗠? We've spent the last 4 years developing a GPU-driven renderer for #n4ce; Rendering in CAD means you don't control the data; you can't spend time optimizing assets like Game Developers do. Engineers are constantly modifying the dataset and innovating. They don't want to wait for long loading bars; They want to start working on their designs and view their surveys instantly. So, what happens when a georeferenced dataset physically exceeds your VRAM budget by 5x and you can't render your model all at once? 𝗧𝗵𝗲 𝗦𝗼𝗹𝘂𝘁𝗶𝗼𝗻: We implemented a mechanism to chop up massive scenes into multiple drawing submits. But that's not as easy as it sounds! There are complex data-dependencies and different memory life-spans for different types of data (e.g., Image Cache vs Raw Geometry Data). Building a GPU-driven renderer capable of this scale meant completely rethinking the pipeline. In the next post, I’ll break down exactly how we moved our curve rendering logic to the GPU, achieving a 𝗺𝗮𝘀𝘀𝗶𝘃𝗲 𝟮𝟱-𝟭𝟬𝟬𝘅 𝗿𝗲𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗶𝗻 𝗺𝗲𝗺𝗼𝗿𝘆 𝘀𝗶𝘇𝗲 while getting infinite zoom polylines in the process. Stay Tuned! We’ve already presented the architecture behind some of our solutions. You can find our presentations and learn more about our work on our website: https://www.devsh.eu/ #CAD #Vulkan #GraphicsProgramming #ComputationalGeometry

  • Karim Sayed and Matt Kielan have been deep in research and made: - Linearly Transformed Cosines Polygonal Lights 2x faster - Arvo 1997 Spherical Triangle Sampling 10x faster than reference implementations - Urena 2013 Solid Rectangle Sampling 20% faster than Arvo 1997 by comparison - Our Urena 2013 implementation is "only" 4x more expensive to draw samples from than a Bilinear distribution or Projected Hemisphere (Lambertian) - Practical Product Sampling via Composing Warps for the above which are an order magnitude at computing MIS weights Meanwhile Kevin Yudi Utama has dusted off our old unpublished invention for sampling environment maps proportionally to their luminosity without having a log2(TexelCount) bandwidth cost. There's enough materials there for 5-6 papers, which we wish we had the time to write, but stay tuned we might write them yet!

Similar pages