Glossary term

Graphics Processing Unit (GPU)

What is a GPU?

A graphics processing unit (GPU) is a specialized processor designed to perform many calculations in parallel, originally to render images, video, and 3D graphics. GPUs now also accelerate other highly parallel workloads, including machine learning, scientific simulation, and video encoding.

In a game or real-time 3D application, the GPU processes scene data to produce the pixels a player sees, rendering the image dozens of times per second.

How does a GPU work?

A GPU gets its speed from parallelism. Instead of relying on a small number of powerful CPU cores to handle tasks sequentially, a GPU divides large workloads into many smaller operations that can be processed in parallel.

The GPU works alongside the CPU, which handles tasks such as game logic and sends instructions to the GPU for graphics and other parallel workloads. For a side-by-side breakdown of the two processors, see CPU vs. GPU.

A GPU relies on several key components and software layers:

  • Cores: Processing units that execute many operations in parallel. Their architecture and naming vary by GPU manufacturer, such as CUDA cores on NVIDIA GPUs and stream processors on AMD GPUs.
  • Video memory (VRAM): Dedicated high-bandwidth memory that stores textures, meshes, frame data, and other resources close to the GPU. Integrated GPUs typically share system memory instead.
  • Shaders: Programs that run on the GPU to perform tasks such as positioning vertices, calculating pixel colors, and processing general-purpose workloads, including those handled by compute shaders.
  • Drivers and graphics APIs: Software layers that allow applications to communicate with the GPU. APIs such as DirectX, Vulkan, Metal, and OpenGL provide interfaces for issuing graphics commands, while drivers translate those commands for specific GPU hardware.

What happens when a GPU renders a frame?

Rendering a single frame follows the graphics pipeline, a sequence of stages that transforms 3D scene data into a 2D image:

  1. Submit: The CPU sends draw calls that specify which meshes, materials, and rendering settings to use.
  2. Process vertices: Vertex shaders transform each mesh's vertices from 3D scene space into positions used for rendering the image.
  3. Rasterize: The GPU converts the resulting triangles into fragments that correspond to potential pixels on the screen.
  4. Shade pixels: Fragment shaders calculate the color of each fragment using information such as textures, lighting, and materials.
  5. Output: The resulting image is written to a frame buffer and presented on the display.

What are the types of GPUs?

GPUs can be grouped into three common types based on where the GPU is located, how it accesses memory, and how its computing resources are delivered:

Where it lives

Integrated GPU

On the same chip or package as the CPU

Discrete GPU
On its own graphics card or module
Virtual (cloud) GPU

In a cloud provider's data center, accessed remotely

Memory

Integrated GPU

Shares system RAM
Discrete GPU
Dedicated VRAM
Virtual (cloud) GPU

Allocated GPU memory on host hardware

Power and heat

Integrated GPU

Generally lower
Discrete GPU
Generally higher and may require dedicated cooling
Virtual (cloud) GPU
Managed by the provider
Typical devices

Integrated GPU

Laptops, phones, tablets, standalone VR headsets
Discrete GPU
Gaming desktops, workstations, gaming laptops
Virtual (cloud) GPU
Cloud servers and virtual machines
Best for

Integrated GPU

Everyday computing, lighter games, and mobile applications
Discrete GPU
High-end games, 3D content creation, and local AI workloads
Virtual (cloud) GPU
AI workloads, remote rendering, and cloud-based applications

Integrated GPUs share system resources with the CPU, which generally makes them more power-efficient and suitable for devices with limited space or power.

Discrete GPUs have their own processing hardware and dedicated memory, providing more graphics and compute performance at the cost of higher power consumption and heat.

Virtual GPUs provide GPU computing resources remotely, allowing applications to use GPU hardware hosted in a data center rather than on the local device.

What's the difference between a GPU and a graphics card?

The terms GPU and graphics card are often used interchangeably, but they refer to different things:

What it is
GPU
The processor that performs the graphics calculations
Graphics card
A hardware component that houses a discrete GPU

What it includes

GPU
Processing cores and on-chip logic
Graphics card
The GPU plus VRAM, cooling, power delivery, and display outputs
Where it's found
GPU
Computers and other devices that process graphics
Graphics card
Desktops and laptops with discrete graphics

Every graphics card contains a GPU, but not every GPU is on a graphics card. Integrated GPUs are built into the same chip or package as the CPU and share system resources rather than using a separate graphics card.

What is a GPU used for?

GPUs are used wherever a workload contains many calculations that can be performed in parallel:

  • Games and real-time 3D: Render frames for games and interactive experiences on phones, PCs, consoles, and other devices.
  • VR and AR: Render immersive views and maintain the high, consistent frame rates needed for comfortable experiences, particularly in VR.
  • AI and machine learning: Accelerate the training and inference of machine learning models, which involve large numbers of parallel mathematical operations.
  • Video and content creation: Accelerate video encoding, rendering, visual effects, and GPU-accelerated workflows in creative applications.
  • Simulation and digital twins: Run physics, particle, lighting, and other computational simulations for engineering, automotive, and industrial visualization.
  • Web graphics: Use APIs such as WebGL to enable browsers to render interactive 2D and 3D graphics with GPU acceleration.

GPU and Unity

In Unity, the GPU renders frames according to the active Universal Render Pipeline (URP) or High Definition Render Pipeline (HDRP) configuration. The GPU Usage Profiler module shows how much GPU time different parts of a frame take, helping developers identify rendering bottlenecks. The GPU Resident Drawer uses GPU instancing to reduce draw calls and CPU overhead when rendering large numbers of objects.

Frequently asked questions (FAQ)

What does it mean when a game is GPU-bound?

A game is GPU-bound when the GPU takes longer to render a frame than the CPU takes to prepare it. Rendering demands such as high resolution, complex shaders, and overdraw can increase GPU frame time and limit the frame rate. Profiling helps identify whether the CPU or GPU is the bottleneck, and Unity's guide to managing GPU usage covers common optimization techniques.

Can a computer run without a GPU?

A computer can run without a dedicated GPU, but most computers that display graphics use an integrated or discrete GPU. Some servers operate without a GPU because they do not need to render graphics, and software rendering can use the CPU to generate images instead. However, CPU-based rendering is generally less efficient than GPU acceleration for demanding real-time 3D graphics.

What is a GPU benchmark?

A GPU benchmark is a test that measures graphics or compute performance using metrics such as frame rate, frame time, or a composite score. Synthetic benchmarks run standardized workloads to compare GPUs, while real-world benchmarks measure performance in actual games or applications. Results can vary depending on resolution, graphics settings, drivers, and the specific workload.

Related terms

3D Rendering

3D Rendering is the process of generating 2D images from 3D scene data using lighting algorithms, texture mapping, and material properties while balancing visual fidelity against performance requirements for real-time applications.

Ray Tracing

Ray tracing is a rendering technique that simulates real light behavior to create physically accurate reflections and shadows.

Frames Per Second (FPS)

Frames per second (FPS) measures how many images a game or application displays each second.