Design software for a cloud native environment
Connect repeatable infrastructure, replaceable processes, durable state, and observable behavior. Assess cloud native design beyond container packaging.
Published by TaigaHow we write
What you will learn
- Distinguish container packaging from cloud native behavior.
- Identify state, retry, and replacement risks in a generated service.
- Define a platform contract that agents and people can verify.
Define the behavior you need
Cloud native practices support repeatable development and operation in public, private, or hybrid environments. CNCF emphasizes systems that remain manageable, observable, and resilient as they change. Containers and orchestration can support this approach. They do not establish all these properties by themselves.
Start with a fictional report service. An AI tool creates an endpoint, a worker, and a container image. A demonstration produces the correct PDF. Before production, the team must answer another question: what happens when the platform replaces the worker during a job?
This is an application design question as well as an infrastructure question. A restart can restore a process while losing its unfinished work.
Separate a process from durable state
The prototype keeps queued jobs and completed reports on the container disk. Replacing the container can remove both. Adding more workers can also produce different answers depending on which worker receives a request.
The revised design uses a durable job store and an approved object store. A request records a job identity. A worker claims the job, creates its result, and records the result location. Access checks still apply when a user downloads the report.
| Concern | Question for the report service |
|---|---|
| State | Which records must survive process replacement? |
| Configuration | How does the same artifact run in each environment? |
| Identity | Which service identity can read the job and write its result? |
| Health | Can the worker accept work, and can it complete that work? |
| Shutdown | What happens to a claimed job when a worker stops? |
| Capacity | Which limit applies first: workers, database, storage, or another service? |
Keep secrets outside the image. Supply them through the approved secret system. Record which configuration changes require a new release or a process restart.
Design retries before adding workers
Suppose the worker saves a PDF and then stops before acknowledging the job. The queue delivers the job again. A second attempt must not create a second customer charge or send conflicting completion messages.
Use an idempotent operation where appropriate. Repeating the same logical request should preserve the intended effect. Define a stable request identity, record the outcome durably, and check what happens at each failure point. AWS describes this technique in its guide to safe retries.
Retries also need limits. Use a timeout, a retry limit, and a delay that avoids simultaneous repeated requests. Preserve failed work for inspection instead of retrying it forever.
Make the desired state reviewable
A declarative configuration states the intended deployment. A controller works to maintain that state. For example, a Kubernetes Deployment manages application replicas and controlled updates. The application must still handle replacement correctly.
Version the infrastructure and application configuration. Review changes through the normal delivery process. Observe actual job completion, queue age, failures, and dependency limits. A running process can still be unable to produce a report.
Choose a platform the team can operate
Cloud native does not require every application to become microservices. A modular application on a managed runtime can meet its requirements. More services introduce more interfaces, deployment decisions, and operational work.
Give the development agent the actual platform contract: supported runtime, identity method, data services, deployment rules, and required evidence. Test interruption and replacement behavior alongside successful requests. Continue with availability and failure boundaries.
Do the exercise
A fictional report service stores jobs and completed files on its container disk. Draw the flow through request, job, file, and download. Mark durable state. Define what happens if the worker stops after writing a file but before confirming the job.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- CNCF: Cloud Native Definition v1.1 ↗
- Kubernetes: Deployments ↗
- AWS Builders’ Library: Making retries safe with idempotent APIs ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.