Shopify and Serverless: Less Ops, More Shops
A lot of our customers have an interesting engineering challenge: 90% of their sales happen in a very short period. What's the best way to handle that? Pay for peak capacity 24/7, or scale everything up ahead of a release?
For us, the answer is serverless. Everyone has a different definition of serverless, but ours starts with one requirement: it must scale to zero. People say serverless is expensive, and it can be if you're running at capacity most of the time. In that case, fine, accept the operational burden of containers. But when a launch brings a sudden spike followed by long periods with little to do, it is a very good fit.
Scaling up and back to zero
With serverless, each order, product or inventory event invokes the capacity it needs. Queues absorb spikes and release work at a rate that Shopify and the connected systems can accept. Concurrency and API limits still need to be designed for, but the platform handles the scaling instead of a cluster we have to operate. It is how our sneaker raffles take millions of entries without falling over.
Most compute costs follow actual use, so a process that runs for a few minutes each day does not need a permanent service. Databases, logs and network traffic still cost money, and continuously busy workloads may be cheaper on dedicated infrastructure. For this layer, however, the people needed to operate that infrastructure often cost more than the compute.
Expect failure, design for recovery
Every external system will be unavailable at some point. An ERP may undergo maintenance, a warehouse API may time out, or Shopify may apply a rate limit. That should not leave somebody restarting a process by hand.
Queues are not exclusive to serverless. Paired with Lambda for processing and DynamoDB for idempotency, though, they provide this recovery without a fleet of workers or a database to operate. A queue holds the work and retries it more slowly while the outage continues. When the service returns, the queue catches up at a controlled rate and the integration recovers without intervention.
Some failures need human input. Retrying will not fix an invalid SKU or a missing address. Those events should be isolated with the reason attached, allowing somebody to resolve them without stopping the rest of the queue.
Solve business problems, not platform problems
Servers need upgrades, security patches, capacity planning, networking, deployment tooling and monitoring. Kubernetes can help manage this, but it also introduces a platform the team must understand and support. That investment is difficult to justify when the job is to validate a product, route an order or update some inventory.
AWS takes on much of that infrastructure work. The development team still owns the application, dependencies, configuration and data. Logs, alarms, deployments and rollbacks also need care. Operations do not disappear, but the team spends more time on business processes and less on the compute platform.
Smaller security boundaries
An integration service can accumulate broad access to products, inventory, orders and external systems. If it is compromised, all those permissions are exposed.
With Lambda, each function can have its own IAM role and only the permissions its task requires. A product validation function does not need permission to cancel orders. Serverless is not secure by default: IAM roles can still be too broad, secrets can leak into logs and dependencies need updating. Smaller boundaries simply reduce the impact when something goes wrong.
An ephemeral environment for every change
QA has never been more important. AI has greatly increased the rate of change, but every change still needs to be shown to be safe.
Serverless and infrastructure as code let us create an isolated environment for each change, test it against real services and remove it afterwards. Functions, queues, permissions and data stores are rebuilt each time, reducing drift and preventing one developer's work from interfering with another's.
This also makes development easier. Cloud services are difficult to reproduce on a laptop, and mocks can diverge from production. Shopify stores and ERP or warehouse test systems may still be shared, but isolating the parts we control makes testing more reliable.
Supporting it with a small team
The main advantage of serverless for us is that a small development and operations team can support a broad integration layer. This depends on consistency: shared deployment, logging, retries, security and monitoring. Each new integration inherits those decisions instead of inventing its own platform. Our testing platform runs on the same pieces: Lambda, S3 and DynamoDB, deployed with one command.
The team can focus on business rules, systems of record, acceptable delays and what happens when systems disagree. Serverless is not right for every workload, but we think it is a strong fit for the integration layer around high-traffic Shopify sites.
If most of your sales land in a few minutes and you want an integration layer that keeps up, then scales back to zero, speak to us at [email protected].