Native Java: when to use it, when not to, and what suits cloud native
By Dorian Chávez · founder of Hábil and integration architect ·
GraalVM, Quarkus, Spring Boot, Java 25's startup cache and CRaC: what each option gains and costs, and which fits your service and your cloud.
In almost every conversation about Java microservices, the same question comes up: should we compile them to native? The promise sounds irresistible: services that start in a fraction of the time and use a fraction of the memory. And it is true, with fine print that almost no one reads: what you gain in startup and memory you pay for in processing capacity, in build time and in operational complexity.
This article is for whoever decides how the Java services of a bank, an insurer, a retail chain or any regulated company are built and run. It answers three questions: when native makes sense and when it doesn't, what to do if your service is new or if it already runs on Spring Boot, and what suits a cloud native architecture.
First, what "native" means
A regular Java service runs on the JVM, the Java virtual machine. It starts, loads its classes and, while it serves requests, an internal compiler (the JIT) keeps optimizing the code that is used most. That is why a Java service takes time to "warm up": it is slow for the first few seconds and very fast afterwards.
Compiling to native —with GraalVM Native Image— does all that work ahead of time, at build. The result is an executable that no longer needs the JVM: it starts much faster and uses much less memory. In exchange, the compiler has to know in advance everything the program can do (this is called the "closed world"), and it loses the ability to keep optimizing with the real load.
Between those two extremes there are intermediate options worth comparing before you decide:
- Java's startup cache (Project Leyden). Java 24 introduced ahead-of-time class loading and linking and Java 25 added ergonomics and profiles; with them the JVM can save to a file the class loading and linking work and the profiles from a training run, and reuse them at startup. The cache depends on the application, the JDK, the operating system and the architecture. It is still the same JVM, with its JIT; it starts faster and, in the Quarkus lab, also used less memory [1][2][3].
- Saving and restoring an already warm process. CRaC, an OpenJDK project available in some Java distributions, and AWS Lambda SnapStart take a "snapshot" of the already started application and restore it. CRaC can restore very quickly; SnapStart can bring Lambda's cold start down to under a second in favorable scenarios (AWS figure; it depends on the application, platform and load), with conditions on state, connections and credentials [4][5].
What the measurements say
The most complete public comparison of the three modes is published by the Quarkus project itself in its official guide, using a test service [6]. It is a vendor's lab with a specific load, so read it as an order of magnitude, not as a promise:
| Mode | Time to first request | Peak throughput (requests per second) | Resident memory (RSS) |
|---|---|---|---|
| Regular JVM | ~4.4 s | ~13,300 | ~304 MiB |
| JVM with startup cache (Leyden) | ~1.9 s | ~12,400 | ~240 MiB |
| Native (GraalVM) | ~0.6 s | ~5,400 | ~95 MiB |
Figures from the Quarkus project's lab (vendor): run of April 21, 2026 with Quarkus 3.34.3, JDK 25.0.2, GraalVM 25.0.2, 4 CPUs and `-Xmx512m`. Resident memory is the RAM the process keeps occupied; peak throughput is the maximum capacity under that load, not that of your production; and the time to first request is neither the deployment time nor the latency your customer feels.
Three takeaways that matter more than the numbers:
- In that lab, native reached the first request about seven times sooner and used about a third of the memory.
- But it processed about 40% of the requests per second that the JVM did. In a long-lived service with constant load, the JVM can deliver more sustained throughput, because its JIT optimizes with what really happens in production; Oracle GraalVM offers profile-guided optimization (PGO) to narrow that gap, but it is not available in the Community edition: confirm edition and license, since this lab did not measure it [6][7]. The result depends on the load, the CPU, the garbage collector and the configuration.
- The startup cache keeps a good part of the advantage at little cost: startup more than twice as fast, almost the same throughput and, here, memory down from 304 to 240 MiB (the effect may vary and the cache adds about 198 MB to the artifact). OpenJDK uses the Spring sample application (PetClinic) as a case study for the cache; no case study replaces measuring your own service [1][3].
And a gap worth stating: we found no independent public measurement that compares the three modes on worst-case latency (the p99: the latency below which 99% of requests fall; it shows the tail, not the absolute maximum). The ones that exist come from vendors, with their own loads. The only serious way to decide is to measure with yours.
The times that matter: build, start, be ready and warm up
When people talk about "fast," four different times get mixed together, and it helps to separate them:
| Time | Regular JVM | JVM with startup cache | Native |
|---|---|---|---|
| Build the service | seconds (about 30 s in the Quarkus example) | seconds, plus a training run to generate the cache | 3 to 10 minutes and 4 to 8 GB of memory [6] (vendor figure; depends on application, platform and load) |
| Start until it serves the first request | ~4.4 s | ~1.9 s | ~0.6 s [6] |
| Be ready to receive traffic | however long it takes to start and connect to its dependencies | same, but it starts sooner | same, but it starts sooner |
| Warm up to its best performance | while the JIT optimizes with the real load | part of the warm-up already comes in the cache [2] | no warm-up: it reaches its performance directly, which in this lab was lower [6] |
Two practical consequences. Your team pays the build time with every change and every patch, not the customer; and "ready" almost never depends only on Java's startup: connecting to the database, the message broker or the identity provider often weighs as much or more.
Kubernetes probes: where startup becomes a real problem
Kubernetes decides whether a service is alive and whether it can receive traffic with three probes [17]:
- Startup (has it finished starting?). Until it passes, the other two are not evaluated. The time allowed for startup is the number of attempts times the interval between them (
failureThreshold × periodSeconds). This probe exists precisely for services that are slow to start. - Liveness (is the process still alive?). If it fails, Kubernetes restarts the container.
- Readiness (can it receive traffic now?). If it fails, Kubernetes stops sending it requests, without restarting it.
With the SmallRye Health extension, Quarkus exposes them at /q/health/live, /q/health/ready and /q/health/started [19]; Spring Boot, at /actuator/health/liveness and /actuator/health/readiness, which it enables automatically when deployed on Kubernetes [20].
An illustrative example for a JVM service that takes about 5 seconds to start:
startupProbe:
httpGet: { path: /q/health/started, port: 8080 }
periodSeconds: 2
failureThreshold: 30 # tolerates up to 60 s of startup
livenessProbe:
httpGet: { path: /q/health/live, port: 8080 }
periodSeconds: 10
readinessProbe:
httpGet: { path: /q/health/ready, port: 8080 }
periodSeconds: 5In native, with half a second of startup, the startup probe hardly waits; with a JVM, it is the one that keeps the cluster from killing the service halfway through coming up. A service that is slow to start doesn't need native to survive deployment: it needs a well-placed startup probe. That probe prevents premature restarts, but it does not shorten the time to readiness: if your scaling target requires being ready sooner, the startup cache or native remain options to measure.
The two mistakes we see most:
- A liveness probe that checks the database or an external service. If that dependency goes down, Kubernetes restarts all the instances at once and turns an external problem into an outage of your own. Spring Boot's documentation warns about this in exactly those terms [20]. Liveness should only answer "the process is alive."
- Having no startup probe and compensating with a fixed delay on liveness. If one day startup takes longer —a loaded node, a slow dependency—, the service enters a restart loop.
Readiness can take dependencies into account, with judgment: if without the database the service can't answer anything useful, it should be taken out of traffic; if it can answer partially, it is better to leave it in and handle the error further up [20].
How much memory a Java service asks for in a container, and why
It is common to see Spring Boot services that request around 500 MB per container, even if their logic is small. This is not a defect of Spring: it is the sum of what the JVM needs to live:
- The heap, where objects live. If it isn't configured, the JVM takes at most a quarter of the memory it detects [22], and inside a container it detects the container's limit [23].
- Loaded classes (metaspace). A framework with many features loads thousands of classes.
- The code already optimized by the JIT (code cache).
- A stack for each thread. A web server with a large thread pool adds up memory even if those threads are waiting.
- Buffers and the garbage collector itself.
That is why what matters in Kubernetes is not the heap size but the total memory of the process (what the system sees as RSS), and the container limit has to cover it with margin. Otherwise Kubernetes kills the container for lack of memory even though the heap looks healthy.
Before going native, there are adjustments that lower memory without changing models: set the percentage of memory the heap takes, size the thread pools to what the service really handles, and choose the collector according to the size of the container. The startup cache, on the other hand, speeds up startup and, in the Quarkus lab, lowered memory from 304 to 240 MiB (~21%); the effect may vary and the cache increases the artifact size [24].
Where native does change the math is density. An illustrative calculation: on a node with 16 GB available for services, about 32 containers of 500 MB fit, or about 160 of 100 MB. If your cloud bill or your on-premises nodes are determined by memory, and you have dozens of small services, that difference pays for native compilation. If you have few large services with constant load, it doesn't.
What native costs and almost no one puts in the proposal
- Building is slow and heavy. The Quarkus guide estimates 3 to 10 minutes and 4 to 8 GB of memory for a native build (vendor figure; depends on application, platform and load), versus a few seconds for the regular one [6]. This repeats with every change, every Java patch and every new dependency, and it shows in the pipeline.
- The "closed world" breaks things silently. What the program discovers at run time —reflection, proxies, serialization, dynamic class loading, resources— has to be declared. Otherwise the binary may compile and fail when running a dynamic path nobody tested; that is why a native pilot has to include integration tests of the binary, not just of the code [8].
- In Spring Boot, when AOT processing or native is enabled, some decisions are frozen at build time. Profiles and properties that change which components are created can no longer be changed at startup; credentials and addresses can [9].
- Diagnosing is different. JFR, the event recording a team uses daily on the JVM, exists in native, but it comes turned off and has to be included at build time [10]; the same goes for other diagnostic tools. Validate beforehand that your operations have what they need to investigate an incident.
- In a regulated company, native adds things to govern: the compiler and its version, the metadata, the component inventory (SBOM) and the binary's tests, all traceable by version. And when a vulnerability appears in a dependency or in the JDK, updating is not enough: you have to rebuild, retest and re-release the binary. It doesn't reduce the security work; it moves it elsewhere.
If your service is new: Quarkus?
For a new service, the right question is not "Quarkus or Spring?" but "what kind of service is it?"
- Quarkus was born for this. It resolves almost everything at build time and its extensions declare whether they are compatible with native [11]; even so, compatibility is checked dependency by dependency. If the service is small, will scale to zero or lives on little memory, native Quarkus can reduce startup and memory when the required extensions are compatible; the Quarkus guide recommends starting on the JVM and moving to native when there is a concrete need. And if it later turns out that the JVM is the better fit, Quarkus also runs very well on it.
- Spring Boot remains a great option when the team already masters it, when the service depends on libraries from its ecosystem or when the logic is complex and long-lived. Since version 3 it officially supports native, and it now also offers Java 25's startup cache [9][12].
- Micronaut and Helidon also offer native paths; compatibility is checked per dependency and use case [13][14].
If your team knows Spring Boot: when does it make sense to move to Quarkus?
It is the most common question, and the honest answer starts with what a comparison doesn't show: the cost of having two frameworks. Two ways to configure, test, monitor and hire. A team that masters Spring Boot is already productive, and that productivity is worth more than a few seconds of startup in a service that runs for weeks without restarting.
Stay on Spring Boot when:
- The service lives a long time with constant load: there the JVM can perform better and startup weighs little; validate it with your traffic profile.
- It depends on Spring ecosystem libraries that have no direct equivalent.
- Java 25's startup cache already solves your timing problem [12].
Moving to Quarkus makes sense when:
- There are new services that will scale to zero, run as functions or live on very little memory, and native does pay for its cost.
- Cluster density is a real cost problem: many small services that don't fit on the nodes.
- The team has room to learn, and you start with a pilot service, not a migration.
What you have to relearn, told from experience. Moving to Quarkus is not changing annotations. It changes the way dependencies are injected (CDI instead of Spring's container), the way data is accessed (Panache, with its own style of entities and repositories, instead of Spring Data), the REST layer, the configuration and the development cycle. We had to relearn a good part of what we took for granted. It is worth it when the type of service calls for it; it is not worth it for everything.
What eases the change, and what doesn't. Quarkus ships a compatibility layer for the most commonly used Spring annotations —dependency injection, web controllers, Spring Data JPA, properties, transactions— so that a Spring team is productive from day one [21]. But it is a bridge, not a copy: it doesn't support @Conditional or @ComponentScan, because Quarkus resolves dependencies at build time, and the guide itself recommends moving over time to the standard CDI annotations [21]. A large Spring service is not "converted" to Quarkus; it is rewritten calmly or it stays where it is.
The rule we use: you don't change frameworks because of fashion or a benchmark. You change when a specific type of service justifies it with numbers, and you start with one.
If you already have Spring Boot: don't jump straight to native
Jumping straight to native with a Spring service that works, "because it starts faster," adds risk when the compatibility of its dependencies hasn't been validated. The order we recommend:
- Upgrade first. Java 25 LTS —or a later version your organization has validated— and the current version of Spring Boot can bring startup and memory improvements; validate it with functional and operational tests.
- Turn on the startup cache. Spring Boot documents it as its recommended option for Java 25 onward [12]. It requires a training and validation flow for the artifact (same application and same Java version). It is a lower-risk change than native and, in many services, it can be enough.
- Evaluate CRaC only if your platform supports it and your team can handle what it demands: closing and reopening connections, refreshing credentials and not leaving secrets inside the snapshot [4][15].
- Go native only with a measured pilot on the service where memory or startup really cost money.
What suits cloud native
"Cloud native" doesn't mean "native." It means designing services that scale, recover and deploy automatically. What suits depends on how each service lives:
| Type of service | What suits | Why |
|---|---|---|
| Functions and services that scale to zero | Native, or SnapStart if you operate on AWS Lambda and validate its restrictions | Startup is paid on every cold request [5][16] |
| Long-lived services with constant load | JVM, with startup cache | The JIT delivers more sustained throughput [6] |
| Unpredictable traffic spikes | Native for the instances added at the peak | They can shorten the time until they are ready; measure it including connections and dependencies |
| Many small services in a cluster | Native if memory is what limits density | More services fit per node |
| Command-line tools and short-lived processes | Native | There is no time to warm up |
In Kubernetes there is a practical detail: a service that is slow to start needs a startup probe with enough time; otherwise the cluster restarts it before it finishes coming up [17]. The startup cache and native reduce that problem; a startup probe that covers the worst-case startup avoids the restarts, although it doesn't turn a slow service into a ready one.
What doesn't hold up
- "Native is always faster." It starts sooner; under sustained load, the JVM usually processes more.
- "Native solves worst-case latency." The tail is still determined by the garbage collector, the network, connection pools and external dependencies.
- "It compiles unchanged." Only if all its dependencies are already prepared for the closed world.
- "Less memory means lower cost." You have to add CPU per request, the pipeline, the artifact size and operations.
- "Leyden already replaces native." Not yet: it keeps the JVM, and ahead-of-time code compilation (JEP 544) is in candidate status, not available in a released JDK version [18].
Where to start
- Classify your services by how they live: functions, long-lived services, spikes, tools.
- Measure before you decide: startup, memory, requests per second and worst-case latency, with your real load.
- Try the cheap things first: upgrade Java and turn on the startup cache.
- Run a native pilot only where startup or memory cost money, with your real dependencies and your pipeline.
- Define the exit criterion before you start: if the pilot doesn't improve the metric that hurts (memory, startup or cost) enough to pay for its complexity, the service stays on the JVM.
- Decide with the numbers in hand, and write down why.
The assessment doesn't start by migrating. We classify the candidate services, measure startup, memory, capacity and worst-case latency with your load, review dependencies and operational controls, and deliver a traceable decision per service: where the JVM, the startup cache, CRaC or native makes sense, what risk remains and what the minimum pilot would be. That way you decide with evidence before committing your platform.
References
- OpenJDK, JEP 483, Ahead-of-Time Class Loading & Linking (JDK 24). https://openjdk.org/jeps/483
- OpenJDK, JEP 514 and JEP 515 (JDK 25). https://openjdk.org/jeps/514 · https://openjdk.org/jeps/515
- OpenJDK, Project Leyden. https://openjdk.org/projects/leyden/
- OpenJDK, CRaC. https://github.com/openjdk/crac
- AWS, Lambda SnapStart. https://docs.aws.amazon.com/lambda/latest/dg/snapstart.html
- Quarkus, guide Building a Native Executable (version 3.40), JVM, AOT cache and native comparison from the project's lab. https://quarkus.io/version/3.40/guides/building-native-image/
- GraalVM, Optimizations and Performance. https://www.graalvm.org/jdk25/reference-manual/native-image/optimizations-and-performance/
- GraalVM, Dynamic Features (reflection, proxies, resources). https://www.graalvm.org/latest/reference-manual/native-image/dynamic-features/
- Spring Boot, Ahead-of-Time Processing and GraalVM Native Images. https://docs.spring.io/spring-boot/reference/packaging/aot.html · https://docs.spring.io/spring-boot/reference/packaging/native-image/introducing-graalvm-native-images.html
- GraalVM, JFR in Native Image. https://www.graalvm.org/latest/reference-manual/native-image/debugging-and-diagnostics/JFR/
- Quarkus, versions and support (3.40 LTS). https://quarkus.io/blog/quarkus-3-40-released/
- Spring Boot, AOT Cache. https://docs.spring.io/spring-boot/reference/packaging/aot-cache.html
- Micronaut, documentation (5.2). https://docs.micronaut.io/5.2.x/core/
- Helidon, Native Image. https://helidon.io/docs/v4/mp/guides/native-image
- Spring Boot, Checkpoint and Restore. https://docs.spring.io/spring-boot/reference/packaging/checkpoint-restore.html
- AWS, Reducing Java cold starts on AWS Lambda functions with SnapStart. https://aws.amazon.com/blogs/compute/reducing-java-cold-starts-on-aws-lambda-functions-with-snapstart/
- Kubernetes, Configure Liveness, Readiness and Startup Probes. https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
- OpenJDK, JEP 544, Ahead-of-Time Code Compilation (proposed). https://openjdk.org/jeps/544
- Quarkus, SmallRye Health. https://quarkus.io/guides/smallrye-health
- Spring Boot, Actuator: Kubernetes Probes. https://docs.spring.io/spring-boot/reference/actuator/endpoints.html
- Quarkus, Quarkus Extension for Spring DI API and Spring compatibility guides. https://quarkus.io/guides/spring-di
- Oracle, Java SE 25 GC Tuning Guide: Ergonomics (default maximum heap, 1/4 of physical memory). https://docs.oracle.com/en/java/javase/25/gctuning/ergonomics.html
- Oracle, The java Command (container detection,
UseContainerSupport). https://docs.oracle.com/en/java/javase/25/docs/specs/man/java.html - Quarkus Performance Lab, JVM, AOT cache and native runs (memory from 304 to 240 MiB with AOT cache); see [6].
- Cloud & infrastructure →
- Modernize the core without slowing the business →
- Release without fear on your own infrastructure →
Does your operation face these challenges?
Prefer email? Write to us at hola@habil.mx