Good architecture often loses to friction.
You know something should probably be a separate service. But the code is already here. The dependencies are already here. The deployment already exists. Adding another handler feels easier than creating another boundary.
So you add it.
Then another one.
Eventually things that made perfect sense together from a business perspective are sharing an operational boundary that makes no sense at all.
We ran into a good example of this recently.
We had a worker processing 16 message types from one SQS queue. Some of them hit a service that performs expensive DB queries with high CPU load. Others made two fast API calls and were done in milliseconds.
Same queue. Same worker pool. Same autoscaling config.
The result was predictable.
The fast messages got stuck behind the slow ones.
Scale up the workers to drain the queue and we increase the number of expensive queries hitting the database in parallel.
Be conservative with scaling to protect the database and the fast messages starve.
There was no good number to put in the autoscaling config because the problem wasn’t the number.
The workloads wanted different scaling policies.

All 16 message types belong to the same domain, by the way.
I think that’s important.
This wasn’t a case where somebody dumped unrelated functionality into the same service. Logically, these messages belong together.
Operationally, they don’t.
Every worker gets an autoscaling definition. Thresholds, step sizes, cooldowns, min/max capacity. It also gets CPU and memory allocation.
You can tune those numbers all day long. If one workload wants low concurrency and more resources while another wants high concurrency and almost no resources, you’re tuning a compromise.
The obvious answer was to separate them.
So that’s what we did.
The fast-path messages got their own queue and worker pool. The heavy path scales conservatively with more resources. The fast path scales aggressively with minimal resources.

Nothing particularly clever here.
The interesting part is that the entire thing took about an hour.
One handler. A few utility files. Two npm dependencies.
We already made services cheap
I’ve been somewhat obsessed with this problem for a long time.
I built my first version of this kind of platform at another company in 2018.
When I came to Hippo, we took the idea much further.
I wrote about the Hippo PaaS back in 2022.
The philosophy behind it was pretty simple: infrastructure should not dictate architecture.
If the correct architecture requires another service, creating that service shouldn’t be a project.
So we built products instead of giving engineers a box of AWS pieces.
Need an HTTP service? That’s a product.
Need a worker? Product.
Need a scheduled runner? Product.
They come with metrics, logging, autoscaling, deployment, networking, secrets, permissions and all the other things a production service needs.
You don’t write Terraform.
You describe what you want.
The platform creates the repository, infrastructure, build configuration, deployment configuration and everything around it.
Most things have sane defaults. In many cases the engineer needs to provide little more than a service name and the team responsible for it.
Back when I wrote about it in 2022, you could create a service and publish it to production in under 20 minutes.
That was four years ago.
Creating a service at Hippo today is not expensive.
We already solved that problem.
And yet people still add things to existing services.
Friction has layers
This is the part I find interesting.
Even after you remove the infrastructure work, the existing service still has gravity.
The file is already there.
The dependencies are already installed.
The tests are there.
The types are there.
You understand the project.
Creating the infrastructure for another worker may be almost free, but you still have to move the code. Figure out which utilities it needs. Bring over dependencies. Fix imports. Get the new project compiling. Run the tests. Fix whatever you missed.
None of this is hard.
That’s almost the problem.
It’s just annoying enough that:
I’ll put it here for now.
feels perfectly reasonable.
Do that enough times and you get what I call a monolithic microservice.
You technically have microservices. Lots of them.
But individual services keep accumulating workloads until they have the same kinds of conflicting operational requirements you were trying to avoid by having service boundaries in the first place.
That’s what happened here.
Sixteen perfectly reasonable message types.
One worker.
One scaling policy.
AI removed another layer
This is where AI coding tools made a real difference for me.
Not because the AI figured out the architecture.
It didn’t.
We looked at the behavior of the system and realized that these workloads couldn’t share a scaling boundary.
That’s the important decision.
Once we made that decision, though, most of what remained was mechanical.
Move this handler.
Bring these utilities.
Add these dependencies.
Fix the imports for the new project structure.
Build it.
Fix what’s broken.
Build it again.
The agent did most of that.
The entire change took about an hour.
That’s where I think a lot of the value of AI coding tools is going to come from.
Not necessarily writing some brilliant algorithm you couldn’t write yourself.
Removing another layer of friction between deciding what the architecture should be and actually making the system look like that.
We spent years making the infrastructure side of this boring.
That was intentional.
I don’t want an engineer deciding whether something should be a service based on how annoying Terraform is, how many dashboards need to be created, or whether they feel like setting up another deployment pipeline.
Those things should already exist.
Now AI is starting to make more of the implementation boring too.
That’s also a good thing.
A new service still isn’t free. It has a lifetime cost. Somebody owns it. It runs in production. It needs to be understood, operated and eventually changed or deleted.
I’m not arguing for turning every function into a service.
I’m arguing that the upfront implementation cost should have as little influence as possible on where we draw an architectural boundary.
In this case the architecture was telling us something very clearly:
These workloads cannot share a scaling policy.
So they shouldn’t share a scaling boundary.
The platform made creating that boundary cheap.
AI made moving the implementation across it cheap.
The part that wasn’t cheap was recognizing that the boundary needed to exist.
That’s the part I want engineers spending their time on.