3DMark
About 3DMark
People talk about a benchmark score from this suite as though there were one number. There is not. 3DMark is a collection of separate benchmarks aimed at different classes of hardware and different rendering approaches, and choosing the wrong one produces a number that means nothing at all.
That is the first thing to understand and the thing most guides skip. A heavy test on a laptop with integrated graphics will crawl and tell you only that the test was too heavy. A light test on a current desktop card will finish instantly and tell you nothing about the card.
Once the right test is running, what you get is a repeatable figure, a comparison against a very large database of other machines, and graphs showing what the hardware did while it worked.
A suite rather than a score
The current lineup covers most of what people own. A heavy rasterised test rendering at four times high definition resolution is the reference for current desktop graphics cards. A lighter version of the same test exists for hardware that cannot manage it, and it runs across other device classes too so a laptop and a handheld can be compared honestly.
Older tests survive because the results database is built on them. The previous mainstream test rendering at an intermediate resolution remains the most widely quoted figure anywhere, and an even older test using the previous generation graphics interface is still the reference for hardware of that vintage.
Ray tracing has its own pair. A heavy current test and an earlier one that established the category, both measuring something the rasterised tests deliberately ignore. Light tests round out the bottom of the range for integrated graphics and mobile hardware.
Which test for which machine
The practical mapping matters more than the names. A current dedicated graphics card belongs in the heavy rasterised test, and if it supports ray tracing then in the heavy ray-tracing test as well, since the two measure different circuitry inside the same card.
A thin laptop or a gaming handheld belongs in the lighter cross-platform test, which is the one designed to produce meaningful numbers on modest hardware and to compare across device categories. Integrated graphics belong in the light test built for them, where they will produce a score rather than a slideshow.
Getting this wrong is the most common mistake with 3DMark, and the symptom is a result that sits absurdly low or high compared with everything else in the database. If you are unsure what you actually have, GPU-Z identifies the card, its clocks and how it is connected, which also catches a card running in the wrong slot.
The processor can invalidate a graphics score
Here is a detail from the documentation that deserves far more attention than it gets. The current heavy test has no processor component at all, deliberately, so that it measures the graphics card alone.
That does not make the processor irrelevant. Its job during the test is building the command lists the card executes, and the documentation states plainly that if it cannot submit work fast enough it becomes the limiting factor and effectively invalidates the graphics result.
So a modest processor paired with a fast card produces a low graphics score, and the low score is not the card’s fault. Anyone benchmarking a new card in an older machine needs to know that before concluding the card was a disappointment.
CPU Profile measures scaling, not speed
The processor test in 3DMark does something more useful than producing a single figure. It runs the same workload at one thread, then two, four, eight, sixteen and finally every thread the processor has, and reports each result separately.
That set of numbers answers a question a single score cannot. A game that uses four threads will not care how many cores you bought, and the profile shows exactly where the curve flattens. Two processors with identical top-end figures can behave very differently at the thread counts games actually use.
For processor work that resembles real rendering rather than game engine simulation, Cinebench approaches the same hardware from a different direction and the two together give a fuller picture.
The stress tests are the most useful part
A single 3DMark run lasts a couple of minutes, which is long enough to measure peak performance and far too short to find anything wrong. The stress test loops for twenty minutes and reports a frame rate consistency percentage alongside graphs of clock speeds and temperatures across the whole run.
That is where real faults appear. A cooling system that cannot sustain load shows as clocks dropping and consistency falling away, which is invisible in a short run and obvious over twenty minutes. Laptops and handhelds in particular perform very differently in minute one and minute eighteen.
Watching the run with MSI Afterburner overlaid adds power draw and fan behaviour to the picture, which helps distinguish a thermal limit from a power limit.
For pushing a card until it actually fails rather than measuring how it holds up, a dedicated graphics stress test is harsher and faster at finding instability.
The stress tests belong to the paid editions, which is the main functional difference between them and the entry version, along with custom resolutions, loop counts and individual settings.
Feature tests and the results database
3DMark also carries small tests measuring individual capabilities rather than overall performance. Ray tracing, mesh shaders, sampler feedback, variable rate shading and the upscaling technologies each get their own, which is how you establish whether a card supports something at a useful speed rather than merely claiming support.
The database is the other half of the product. Tens of millions of results are searchable, and 3DMark compares your run against machines with similar hardware automatically. Results have to come from a current build to be included, so an old installation produces figures that will not appear in comparisons.
Treat the comparison charts with scepticism though. That database is heavily populated by enthusiasts running overclocked hardware in configurations nobody would live with, so sitting below average is normal and means considerably less than the interface implies.
What a score does not tell you
The honest limitation of every synthetic benchmark applies here. These tests measure the tests. They use the same techniques current games use, which is why the numbers correlate reasonably well with gaming performance, and correlation is not prediction.
Where a score earns its keep is comparison against itself. Run it before and after changing a setting, a driver, a thermal paste application or a case fan arrangement, and the difference is real information. Run it to find out whether a particular game will manage sixty frames a second and you are guessing with extra steps.
The other honest use is spotting hardware that is underperforming its class. A card scoring well below others of the same model is telling you something is wrong, whether that is thermal, a power connection, a driver or a slot running at reduced width. Confirming stability afterwards is a separate job, and a stability tester that hunts errors rather than measuring speed is the tool for that half.
Conclusion
3DMark is the right tool when the question is comparative. Confirming a new card performs like others of its model, measuring what a cooling change actually achieved, finding out whether a laptop throttles after fifteen minutes, or seeing where a processor stops benefiting from more threads are all questions it answers precisely and repeatably.
Two things stop it being the whole answer. A score is not a frame rate in the game you care about, whatever the comparison charts imply, and the database you are being measured against is full of machines tuned by people who enjoy tuning machines. Pick the test that matches your hardware, use the stress test rather than the single run when something feels wrong, and treat your own earlier results as the comparison that matters most.
Pros & Cons
- Separate tests for different hardware classes, so results stay meaningful across the range
- The heavy current test excludes the processor deliberately, isolating graphics performance
- CPU Profile reports scaling across thread counts rather than one aggregate figure
- Stress tests report frame rate consistency plus clock and temperature graphs over twenty minutes
- Feature tests establish whether individual capabilities work at usable speed
- Cross-platform tests allow a laptop, a handheld and a phone to be compared directly
- Results database holds tens of millions of entries for automatic comparison
- Choosing the wrong test for your hardware produces a meaningless number
- A modest processor can invalidate a graphics score without the interface saying so
- Stress tests and custom settings belong to the paid editions
- The comparison database is skewed by overclocked and unrepresentative systems
- Results from older builds are excluded from comparisons entirely
- Synthetic scores correlate with gaming performance without predicting it
Frequently asked questions
The heavy rasterised test for a current dedicated graphics card, the heavy ray-tracing test as well if the card supports it, the lighter cross-platform test for a laptop or handheld, and the light test for integrated graphics. Running a test that is too heavy for the hardware wastes the result.
Frequently the processor. It builds the command lists the card executes, and if it cannot keep up it caps the graphics score without any warning appearing. Thermal limits, a reduced-width slot and driver problems produce the same symptom.
No, and it does not claim to. It uses the same rendering techniques games use, so the numbers correlate, but a score is a comparison figure rather than a forecast for any particular title.
Consistency rather than peak speed. It loops for twenty minutes and reports how steady the frame rate stayed, with graphs of clock speeds and temperatures, which is how thermal throttling and cooling faults reveal themselves.
How performance scales with thread count, measured separately at one, two, four, eight, sixteen and maximum threads. That reveals where a processor stops gaining from additional threads, which matters because games use far fewer than modern processors provide.
No. Stress testing, custom resolutions, loop counts and individual test settings all belong to the paid editions, while the entry version runs the standard benchmarks at their default settings.


(9 votes, average: 3.78 out of 5)