Parallel streams — when do they actually help?
Hard.parallel() splits the source and runs chunks on the common ForkJoinPool. It pays off for large, easily split data (arrays, ArrayList, ranges) with CPU-heavy, independent work per element. It hurts for small inputs, blocking I/O, LinkedList sources, ordering and shared state.
How it works
- Split. The stream asks the source's
SpliteratortotrySplit()again and again, building a tree of chunks. - Run. Each chunk becomes a task on
ForkJoinPool.commonPool(). That pool hasavailableProcessors() - 1workers by default, and the calling thread helps too. - Combine. Partial results are merged up the tree with the combiner (
reduce,collect). - The bill. Splitting, task scheduling, and merging all cost time. Parallel only wins when the per-element work across all elements is big enough to dwarf that overhead.
What helps:
- Lots of elements, or few elements that are each expensive (pure CPU work).
- Sources that split by index: arrays,
ArrayList,IntStream.range. - Stateless operations and cheap, associative merges (
sum,max,count).
What hurts:
- Small collections. Overhead beats the gain.
- Blocking I/O inside the pipeline. It ties up the shared pool for the whole JVM.
LinkedList,Stream.iterate,BufferedReader.lines(): they split badly.- Order-sensitive steps on ordered streams:
limit,findFirst,sorted,forEachOrdered. - Boxing (
Stream<Integer>instead ofIntStream) and merges that copy a lot (toListof big chunks, grouping into maps).
Example
static boolean isPrime(int n) {
if (n < 2) return false;
for (int d = 2; (long) d * d <= n; d++) if (n % d == 0) return false;
return true;
}
// Good fit: big range, CPU-only work, cheap merge
long primes = IntStream.range(0, 5_000_000).parallel().filter(Main::isPrime).count();
// Broken: many threads adding to a plain ArrayList
List<Integer> found = new ArrayList<>();
IntStream.range(0, 10_000).parallel().filter(Main::isPrime).forEach(found::add); // lost items or exceptions
// Correct: let the stream build the list
List<Integer> safe = IntStream.range(0, 10_000).parallel().filter(Main::isPrime).boxed().toList();Measure with JMH or at least a warmed-up loop before claiming a speedup.
Edge cases
reduce(identity, op)must use a true identity.reduce(10, Integer::sum)adds 10 once per chunk in parallel, so the answer depends on how the data split.findAny()andunordered()let a parallel stream skip the bookkeeping that keeps encounter order.- Running a parallel stream inside
myPool.submit(...)makes it usemyPool. This works in practice but is an implementation detail, not a documented guarantee. - The common pool's size can be set with
-Djava.util.concurrent.ForkJoinPool.common.parallelism=N.
Common mistakes
- Adding
.parallel()to everything as a free speed boost. - Calling a remote service or database in a parallel stream. For blocking work, use an executor (virtual threads on Java 21+).
- Mutating shared state from lambdas instead of using
collect/reduce.
Likely follow-up
"Why is parallel().sorted().forEachOrdered() often no faster than sequential?" Sorting needs every element before it can emit any, and forEachOrdered forces results back into a single ordered line, so most of the parallel work gets serialized again.
Get every deep dive in the app
Coming soon to the App StoreComing soon to Google Play