August 25, 2026
Spring Boot on AWS ECS: The Deployment Issues Nobody Warns You About
A few months back, I moved a Spring Boot microservice from a single EC2 box to ECS Fargate. On paper it was a two-day task. “It’s just a…

By Prince kumar Maurya
8 min read
A few months back, I moved a Spring Boot microservice from a single EC2 box to ECS Fargate. On paper it was a two-day task. "It's just a container," I told my manager. "Should be smooth."
It took eleven days. Not because ECS is bad — it's actually pretty solid once you understand its quirks — but because every single assumption I carried over from running Spring Boot on a plain VM turned out to be wrong in some small, annoying way. This post is everything I wish someone had told me before I started, written from the other side of that mess.
If you're planning your first (or fifth, still-painful) Spring Boot deployment on ECS, save yourself some 2 AM Slack messages to your team.
The setup, for context
Nothing exotic: a Spring Boot 3.2 service, Java 21, built with Maven, packaged into a Docker image, deployed to ECS Fargate behind an Application Load Balancer, with an RDS Postgres instance and a Redis cache sitting in the same VPC. Standard stuff. This is exactly why the problems were so frustrating — there was no unusual architecture to blame.
Issue #1: The container gets OOM-killed and Spring Boot never even gets a chance to log why
This was the first wall I hit. I set the Fargate task to 512 MB memory (seemed reasonable for a simple REST service), deployed, and within about 90 seconds the task died. No stack trace. No graceful shutdown log. Just gone, replaced by ECS, dies again, replaced again — the classic crash loop.
Checked CloudWatch Logs. Nothing useful — the JVM didn't even get to print its startup banner half the time.
Turns out the issue was container memory versus JVM heap. By default (this got better in newer JDKs, but it still bites people), the JVM doesn't automatically understand cgroup memory limits the way you'd expect, and even when it does respect -XX:MaxRAMPercentage, Spring Boot's own startup process — class loading, Spring context initialization, Tomcat thread pools — can spike memory well past what a 512 MB container can survive during startup, even if steady-state usage is fine.
The fix that actually worked:
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 -XX:InitialRAMPercentage=50.0"ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 -XX:InitialRAMPercentage=50.0"And bumping the task to 1024 MB. Not because the app needs a gigabyte to run — it doesn't, steady state is around 280 MB — but because startup is the expensive part, and Fargate doesn't let you burst memory the way a shared EC2 host sometimes does.
Lesson: size your ECS task memory for your worst moment (startup, GC pause under load), not your average moment. I now always test with docker stats locally under a simulated load before I ever touch the ECS task definition.
Issue #2: Health checks that lie to you
ECS uses two separate health check mechanisms if you're behind an ALB — the ALB target group health check, and (optionally) the container health check in the task definition. I only configured one. Guess which one caused the outage.
I had /actuator/health wired up as the ALB target group check, hitting it every 15 seconds. That part worked fine. What I didn't realize is that Spring Boot's default health indicator for a database will happily report UP even while a connection pool is exhausted and every actual request is timing out, because the health check just grabs a connection to run SELECT 1 — it doesn't check if the pool has room for real traffic.
So during a traffic spike, real users were getting 504s from the ALB while the health check kept passing, kept traffic flowing to a task that was functionally dead. ECS had no reason to replace it.
What fixed it wasn't a code change — it was actually tuning HikariCP so the pool would surface exhaustion faster and adding a custom health indicator that checks pool utilization, not just connectivity:
@Component
public class ConnectionPoolHealthIndicator implements HealthIndicator {
private final HikariDataSource dataSource;
@Override
public Health health() {
HikariPoolMXBean pool = dataSource.getHikariPoolMXBean();
int active = pool.getActiveConnections();
int total = dataSource.getMaximumPoolSize();
if (active >= total) {
return Health.down()
.withDetail("activeConnections", active)
.withDetail("maxPoolSize", total)
.build();
}
return Health.up().build();
}
}@Component
public class ConnectionPoolHealthIndicator implements HealthIndicator {
private final HikariDataSource dataSource;
@Override
public Health health() {
HikariPoolMXBean pool = dataSource.getHikariPoolMXBean();
int active = pool.getActiveConnections();
int total = dataSource.getMaximumPoolSize();
if (active >= total) {
return Health.down()
.withDetail("activeConnections", active)
.withDetail("maxPoolSize", total)
.build();
}
return Health.up().build();
}
}It's not a perfect solution, but it's a much more honest one. Health checks that only check "is the process alive" are close to useless in production. You want a health check that reflects whether the service can actually do its job right now.
Issue #3: Task definition environment variables silently not updating
This one cost me about four hours of pure confusion, and I still feel a little dumb about it.
I updated an environment variable in the task definition — bumped a feature flag value — registered a new task definition revision, updated the service. Deployment showed as successful in the console. New tasks came up healthy. And yet the app was still behaving like the old value was set.
The problem: I'd updated the task definition, but the ECS service was still pointed at the previous task definition revision because I'd used aws ecs update-service without explicitly passing --task-definition, assuming (wrongly) that "force new deployment" would pick up the latest revision automatically. It doesn't. --force-new-deployment redeploys the currently configured revision, it does not bump you to the latest one.
# This does NOT pick up your new task definition revision
aws ecs update-service --cluster my-cluster --service my-service --force-new-deployment
# This does
aws ecs update-service --cluster my-cluster --service my-service \
--task-definition my-task-def:47# This does NOT pick up your new task definition revision
aws ecs update-service --cluster my-cluster --service my-service --force-new-deployment
# This does
aws ecs update-service --cluster my-cluster --service my-service \
--task-definition my-task-def:47Small thing. Obvious in hindsight. But it's exactly the kind of gap between "the deploy pipeline says success" and "the thing you wanted actually happened" that makes ECS debugging feel harder than it should.
Issue #4: Graceful shutdown, or the lack of it
Every time I did a rolling deployment, I'd see a small burst of 502s right at the moment old tasks got drained. Not huge — maybe a dozen requests over ten seconds — but enough that our error-rate alert would fire on every single deploy, which meant the team started ignoring that alert, which is exactly how you end up missing a real incident later.
The root cause: ECS sends a SIGTERM to the container, then waits stopTimeout seconds (default 30, I had it at 30) before sending SIGKILL. Spring Boot, by default, does handle SIGTERM and starts shutting down — but the ALB doesn't know a task is being drained the instant SIGTERM fires. There's a window where the ALB can still route a request to a task that's already begun shutting down its Tomcat connector.
Two changes fixed most of it:
- Enabled
server.shutdown=gracefuland set aspring.lifecycle.timeout-per-shutdown-phaseso in-flight requests get to finish:
server:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 20sserver:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 20s- Added a deregistration delay on the ALB target group so it stops sending new requests to a task before ECS actually kills it:
Target group attribute: deregistration_delay.timeout_seconds = 25Target group attribute: deregistration_delay.timeout_seconds = 25The two numbers need to work together — your app's graceful shutdown window should roughly match or be a bit shorter than the ALB's deregistration delay, otherwise you're either killing the app mid-request or letting it sit idle waiting for a load balancer that already gave up on it.
This didn't get us to zero errors during deploy — nothing ever really does — but it took us from a dozen 502s per deploy to basically none over a two-week period of watching it closely.
Issue #5: Cold starts and the ALB timing out before Spring Boot is actually ready
This is subtle and it only shows up under specific conditions: fast auto-scaling events, or deployments where several tasks start simultaneously.
Spring Boot with a decent number of beans, some Feign clients, Liquibase migrations on startup — that can take 12–20 seconds to reach a state where it's actually ready to serve traffic, even though the container itself started in under a second. If your ALB health check interval and healthy threshold are too aggressive, or if your ECS healthCheckGracePeriodSeconds is too low, ECS can mark a task as unhealthy and cycle it before Spring Boot ever finishes booting — leading to a task that never gets a fair chance, forever restarting.
We were running healthCheckGracePeriodSeconds: 30 for a service that, under cold JVM start plus Liquibase migrations, sometimes took 35-40 seconds to become ready. Bumping that grace period to 60 fixed roughly 90% of our "flapping task" incidents.
{
"healthCheckGracePeriodSeconds": 60
}{
"healthCheckGracePeriodSeconds": 60
}I'd also recommend splitting /actuator/health/liveness and /actuator/health/readiness if you're on Spring Boot 2.3+, since Kubernetes-style probes map cleanly onto this, and even without Kubernetes, having a distinct readiness signal helps you reason about "is this container alive" versus "is this container ready for traffic" as two genuinely different questions.
Issue #6: Logs disappearing when a task dies mid-crash
When a task OOM-kills or crashes hard, there's a real chance the last few log lines never make it to CloudWatch, because the log driver needs a moment to flush and a hard kill doesn't give it one. This makes debugging crash loops maddening — you get 200 lines of totally normal startup logs and then just… nothing, right when you need the error most.
A few things that helped:
- Switching from the default
awslogsbuffering to make sure log flush intervals were short - Adding a dedicated uncaught exception handler that logs synchronously before rethrowing, for anything happening during startup
- Turning on ECS Exec (
aws ecs execute-command) for the dev/staging cluster so I could actually shell into a running (or about-to-die) task and look at the JVM directly withjstat/jcmd, instead of purely relying on logs after the fact
That last one turned out to be the single highest-leverage change of the whole project. Being able to run aws ecs execute-command --cluster staging --task <task-id> --container app --interactive --command "/bin/sh" and poke around a live container beats guessing from log fragments every time.
Issue #7: IAM task role permissions that "should" work but don't
Not a Spring Boot issue specifically, but it ate a whole afternoon so it's going in here. The service needed to read a secret from AWS Secrets Manager for the database password. I attached a policy to the task role, deployed, and got:
software.amazon.awssdk.services.secretsmanager.model.SecretsManagerException:
User: arn:aws:sts::...:assumed-role/.../ecs-task is not authorized to perform: secretsmanager:GetSecretValuesoftware.amazon.awssdk.services.secretsmanager.model.SecretsManagerException:
User: arn:aws:sts::...:assumed-role/.../ecs-task is not authorized to perform: secretsmanager:GetSecretValueExcept the policy clearly allowed secretsmanager:GetSecretValue on that exact ARN. The mistake: I'd attached the policy to the task execution role, not the task role. These are two different IAM roles in ECS and it is genuinely one of the more confusing parts of the whole platform if you're new to it.
- The execution role is what ECS itself uses to pull the container image, fetch secrets referenced in the task definition, and write logs.
- The task role is what your application code assumes when it makes AWS SDK calls at runtime.
Since I was fetching the secret manually inside the Spring Boot app (rather than injecting it as a task definition secret), I needed the task role to have that permission, not the execution role. Once I fixed which role got which policy, it worked instantly.
What I'd tell someone starting this today
If I were setting this up again from scratch, in order:
- Set memory generously for startup, tune it down later once you've measured steady state with
docker statsor CloudWatch container insights. - Write a health check that reflects actual readiness to serve traffic, not just "the JVM is running."
- Always deploy with an explicit task definition revision, never rely on force-new-deployment alone.
- Configure graceful shutdown in Spring Boot and ALB deregistration delay together, not just one of them.
- Set
healthCheckGracePeriodSecondsbased on your actual cold-start time, measured, not guessed. - Turn on ECS Exec in non-prod from day one. You'll use it more than you expect.
- Double check which IAM role — execution vs task — actually needs the permission you're granting.
None of these are individually hard. What made the whole thing painful for me was that each one hides behind a slightly misleading symptom — a crash that looks like a bug, a health check that looks fine, a deploy that looks successful. ECS mostly does exactly what you tell it to do. The trouble is figuring out what you actually told it, versus what you thought you told it.
If you're mid-migration right now and something is quietly on fire, I hope one of these seven is your problem. It usually is.