Skip to content

Audit of Cloudemu #1309

Description

@NitinKumar004

CloudEmu — Comprehensive Gaps, Issues & Product Improvements

Overview

CloudEmu already has a strong technical foundation for multi-cloud simulation across AWS, Azure, and GCP. It provides cloud APIs, Kubernetes APIs, resource discovery, networking/topology simulation, cost simulation, monitoring, chaos/fault injection, an in-memory Go API, and SDK-compatible HTTP endpoints.

The biggest opportunity is now not simply adding more cloud APIs.

The next stage should focus on making CloudEmu a complete, reliable, developer-friendly cloud simulation platform.

The goal should be:

Make CloudEmu as easy to start and use as Minikube, while providing the capabilities of a programmable multi-cloud simulation environment.


1. Persistence

Current problem

CloudEmu is primarily in-memory.

A typical workflow is:

Start CloudEmu
    ↓
Create resources
    ↓
Run tests
    ↓
Stop CloudEmu
    ↓
Resources disappear

This is excellent for isolated tests, but inconvenient for development.

Users may want to create an environment once and continue using it tomorrow.

What should be added

Support persistent state:

cloudemu/
└── state/
    ├── resources
    ├── networking
    ├── kubernetes
    ├── IAM
    └── metadata

Possible storage modes:

memory
file
SQLite
embedded database
custom backend

Example:

cloudemu start --persist ./cloudemu-data

After restarting:

cloudemu start --persist ./cloudemu-data

The previous environment should be restored.


2. Snapshots

Persistence alone is not enough.

Users should be able to capture an entire simulated cloud environment.

Example:

cloudemu snapshot create dev-environment

Then:

cloudemu snapshot list

And:

cloudemu snapshot restore dev-environment

Potential snapshot contents:

AWS resources
Azure resources
GCP resources
Kubernetes resources
network topology
security rules
IAM policies
metrics
alarms
cost state
simulation configuration
fake clock

This would make CloudEmu extremely useful for reproducible integration tests.


3. Time Travel / Deterministic Clock

CloudEmu already has the concept of a fake clock.

This should become a first-class developer feature.

Example:

cloudemu time now
cloudemu time advance 24h
cloudemu time advance 30d

Then developers could test:

resource creation
        ↓
24 hours
        ↓
metrics
        ↓
billing
        ↓
alarm
        ↓
scheduled event

without actually waiting.

An even more powerful feature would be:

cloudemu snapshot create before-test
cloudemu time advance 30d
cloudemu snapshot restore before-test

This would make time-dependent cloud testing deterministic.


4. Proper Lifecycle CLI

The CLI should feel like a real developer tool.

Instead of only starting the server, provide:

cloudemu start
cloudemu stop
cloudemu restart
cloudemu status
cloudemu logs
cloudemu version
cloudemu doctor

For example:

$ cloudemu status

CloudEmu
──────────────
Status:     running
Version:    x.x.x
Endpoint:   localhost:4566

Providers:
  AWS       ✓
  Azure     ✓
  GCP       ✓

Resources:
  AWS       42
  Azure     18
  GCP       31
  K8s       24

The goal should be a UX similar to:

minikube
kind
localstack
docker compose

5. Profiles / Isolated Environments

Developers should be able to maintain multiple CloudEmu environments.

Example:

cloudemu profile create development
cloudemu profile create testing
cloudemu profile create demo

Then:

cloudemu profile use development

Each profile should have independent:

resources
state
network
Kubernetes
IAM
metrics
alarms
configuration

This would allow:

developer environment
       +
CI environment
       +
demo environment

without conflicts.


6. Environment Files

CloudEmu should support declarative environments.

For example:

name: ecommerce

providers:
  aws:
    region: ap-south-1

resources:
  - type: ec2
    name: app-server

  - type: rds
    name: database

  - type: s3
    name: assets

  - type: eks
    name: production

Then:

cloudemu apply cloudemu.yaml

and:

cloudemu destroy cloudemu.yaml

This would make CloudEmu much more useful for reproducible integration tests.


7. Terraform Integration

CloudEmu should ideally become a first-class Terraform testing target.

Example:

Terraform
    ↓
CloudEmu
    ↓
AWS / Azure / GCP simulation

Developers could run:

terraform apply

without creating real cloud resources.

This would allow testing:

Terraform modules
provider configuration
resource dependencies
networking
IAM
Kubernetes

before touching a real cloud account.


8. Pulumi / Infrastructure-as-Code Support

Similar support should be considered for:

Pulumi
CDK
Crossplane
OpenTofu

The broader goal should be:

Any infrastructure-as-code tool that talks to supported cloud APIs should be able to target CloudEmu.


9. Better Web Dashboard

CloudEmu already has a powerful underlying resource model, but a developer-friendly dashboard would make that capability much easier to understand.

A dashboard could provide:

CloudEmu Dashboard
──────────────────────────────

AWS       42 resources
Azure     31 resources
GCP       27 resources

Kubernetes
  EKS     2 clusters
  AKS     1 cluster
  GKE     2 clusters

And resource views:

EC2
├── instance
├── status
├── CPU
├── network
├── tags
└── cost

10. Topology Visualization

The networking engine is one of CloudEmu's most interesting capabilities.

The dashboard should visualize:

VPC
 │
 ├── Subnet
 │     ├── EC2
 │     └── EKS
 │
 ├── Route Table
 │
 ├── Security Group
 │
 └── RDS

Users should be able to visually understand:

Who can connect to whom?
Why is this connection blocked?
Which route is being used?
Which security rule allowed/blocked it?

This would make the topology engine much more accessible.


11. Network Debugging

Expose network simulation through CLI/API.

Example:

cloudemu network connect-check \
  --source ec2/app \
  --destination rds/database \
  --port 5432

Output:

Source:
  ec2/app

Destination:
  rds/database:5432

Result:
  BLOCKED

Reason:
  SecurityGroup rule denied traffic

Also provide:

cloudemu network trace ...
cloudemu network explain ...

This could become one of CloudEmu's signature features.


12. Better Cost Simulation

Cost simulation should become more visible.

Example:

cloudemu cost estimate

Output:

AWS
────────────────
EC2              $42.10
RDS              $31.20
S3                $2.40

Azure
────────────────
VM                $28.40
Storage            $1.20

GCP
────────────────
GKE               $35.20

Estimated total:
$140.50/month

Also provide:

cloudemu cost forecast
cloudemu cost compare
cloudemu cost explain

This could eventually allow:

AWS architecture
       ↓
estimated cost

Azure equivalent
       ↓
estimated cost

GCP equivalent
       ↓
estimated cost

13. Cross-Cloud Cost Comparison

Because CloudEmu supports AWS, Azure, and GCP, it could provide a unique capability:

Application workload
        ↓
┌───────┼────────┐
AWS    Azure     GCP
 ↓       ↓        ↓
cost    cost     cost

Developers could compare simulated infrastructure costs before deployment.

This is an area where CloudEmu could differentiate itself from traditional cloud emulators.


14. Stronger Chaos Engineering

Chaos should become a first-class feature.

Examples:

cloudemu chaos latency \
  --service ec2 \
  --delay 500ms
cloudemu chaos error \
  --service s3 \
  --rate 20%
cloudemu chaos throttle \
  --service dynamodb
cloudemu chaos network \
  --source app \
  --destination database \
  --action deny

Potential scenarios:

latency
timeout
throttling
5xx
connection reset
network partition
DNS failure
resource unavailable
rate limiting
partial failure

15. Chaos Scenarios

Instead of configuring individual failures manually, support reusable scenarios.

Example:

scenario: database-failure

steps:
  - inject:
      target: rds/database
      type: latency
      value: 2s

  - inject:
      target: rds/database
      type: error
      rate: 30%

  - wait: 60s

  - recover:
      target: rds/database

Then:

cloudemu chaos run database-failure.yaml

This would make reliability testing much easier.


16. Kubernetes — Continue Expanding the Simulation Layer

CloudEmu already has significant Kubernetes API support across:

EKS
AKS
GKE

The next goal should be deeper Kubernetes simulation where useful.

Potential areas:

Deployment
ReplicaSet
StatefulSet
DaemonSet
Job
CronJob
Service
Ingress
ConfigMap
Secret
ServiceAccount
RBAC
PV
PVC
NetworkPolicy

However, CloudEmu does not necessarily need to become a full Kubernetes runtime.

The distinction should remain:

CloudEmu
    =
Kubernetes API + cloud control-plane simulation

rather than:

CloudEmu
    =
another Kubernetes distribution

17. Kubernetes Controller Simulation

A useful middle ground would be lightweight controller simulation.

For example:

Deployment
   ↓
ReplicaSet
   ↓
Pods

without actually running containers.

This would allow testing cloud-management software that depends on Kubernetes state transitions.


18. IAM Enforcement

CloudEmu already has IAM policy evaluation capabilities.

The next step should be request-level enforcement.

Example:

Application
    ↓
API request
    ↓
IAM evaluator
    ↓
Allowed / Denied

Test:

User A → S3 GetObject → ALLOWED

User A → S3 DeleteBucket → DENIED

This should eventually support:

identity
policy
resource policy
conditions
actions
resources
effect

and return useful denial explanations.


19. Cloud Authentication Simulation

Eventually support simulation of cloud authentication mechanisms where practical:

AWS signing
Azure authentication
GCP authentication
temporary credentials
token expiration
credential failure

This would help test applications that don't simply assume an already-authenticated client.


20. Resource Dependency Graph

CloudEmu should understand relationships between resources.

Example:

EKS
 │
 ├── VPC
 │
 ├── Subnets
 │
 ├── Security Groups
 │
 └── IAM

The system should be able to answer:

What depends on this resource?

What will break if I delete it?

What resources were created because of this resource?

CLI:

cloudemu resource graph ec2/app

21. Better Resource Discovery

The existing discovery capabilities could become a unified query system.

Example:

cloudemu resources list

Filters:

cloudemu resources list --provider aws
cloudemu resources list --type compute
cloudemu resources list --tag environment=dev
cloudemu resources list --region ap-south-1

And eventually:

cloudemu query \
  "all compute resources connected to database resources"

22. Seed Data / Fixtures

CloudEmu should provide reusable environment fixtures.

Example:

cloudemu seed ecommerce
cloudemu seed microservices
cloudemu seed kubernetes
cloudemu seed multi-cloud

A fixture could create:

VPC
subnets
security groups
EC2
RDS
S3
EKS
IAM
monitoring

This would dramatically simplify demos and integration tests.


23. Testcontainers / CI Experience

CloudEmu should become extremely easy to use in CI.

Desired experience:

CI
 │
 ├── Start CloudEmu
 │
 ├── Apply environment
 │
 ├── Run application tests
 │
 ├── Run chaos tests
 │
 └── Destroy environment

Ideally:

cloudemu.Start()
defer cloudemu.Stop()

or a Testcontainers module that makes this automatic.


24. Parallel Test Isolation

A major testing requirement is:

Test A → environment A
Test B → environment B
Test C → environment C

without resources leaking between tests.

CloudEmu should support:

isolated environment
isolated state
isolated network
isolated Kubernetes
isolated clock
isolated metrics

This is especially important for parallel Go tests.


25. Recording / Replay

CloudEmu could record cloud interactions:

Application
    ↓
CloudEmu
    ↓
record requests/responses

Then:

cloudemu replay recording.json

This would allow deterministic reproduction of complicated cloud interactions.


26. Better Observability

CloudEmu itself should expose:

request count
latency
errors
resource operations
chaos events
network decisions
IAM decisions
cost changes

Potential endpoints:

/metrics
/events
/audit
/traces

This is especially useful when debugging integration tests.


27. Audit Log

Every important simulation event could produce an audit event:

{
  "provider": "aws",
  "service": "ec2",
  "operation": "RunInstances",
  "resource": "app-server",
  "timestamp": "...",
  "result": "success"
}

This would make CloudEmu much easier to debug.


28. Documentation / Capability Discovery

One of the biggest current usability problems is that CloudEmu has many capabilities distributed throughout the repository.

A new user may see:

AWS + Azure + GCP emulator

and completely miss:

Kubernetes
topology
cost
chaos
metrics
IAM
resource discovery
fake clock

Create one authoritative capability manifest:

providers:
  aws:
    services:
      ec2:
        operations: [...]
      s3:
        operations: [...]

  azure:
    ...

  gcp:
    ...

Then generate:

README
documentation
CLI help
API reference
coverage matrix

from that manifest.

This also prevents documentation from becoming stale.


29. Automated Capability Tests

Every supported operation should have an automated test.

For example:

operation
   ↓
implementation
   ↓
SDK test
   ↓
HTTP test
   ↓
Go API test

This prevents documentation from claiming support that doesn't actually work.


30. Standardized Service Structure

The current architecture is strong, but naming and directory conventions should be standardized.

Every service should follow the same structure:

service/
├── api/
├── model/
├── driver/
├── mock/
├── wire/
├── tests/
└── docs/

Or another clearly defined convention.

The important thing is consistency.

A contributor should be able to add a new service without first reverse-engineering the repository.


31. Provider Conformance Tests

AWS, Azure and GCP implementations should share a common conformance philosophy.

Example:

Storage interface
      ↓
AWS implementation
Azure implementation
GCP implementation

Then common behavior can be tested consistently.

This is particularly important as CloudEmu grows.


32. Error Fidelity

CloudEmu should reproduce realistic cloud errors.

For example:

ResourceNotFound
AccessDenied
InvalidParameter
Conflict
Throttling
LimitExceeded
DependencyViolation
Unauthorized

Errors should ideally match:

HTTP status
error code
message structure
request ID
provider-specific format

This allows applications to test real error-handling logic.


33. Rate Limiting

Cloud APIs have rate limits.

CloudEmu should allow configurable limits:

rate_limits:
  ec2:
    requests_per_second: 20

  s3:
    requests_per_second: 100

Then applications can test:

throttling
retry
backoff
queueing
circuit breakers

34. Multi-Region Simulation

CloudEmu should eventually model:

AWS
 ├── ap-south-1
 ├── us-east-1
 └── eu-west-1

Azure
 ├── centralindia
 └── eastus

GCP
 ├── asia-south1
 └── us-central1

Then simulate:

region failure
latency
cross-region networking
replication
regional cost

This would greatly increase the usefulness of the simulation engine.


35. Multi-Account / Multi-Subscription / Multi-Project

Real cloud environments have organizational boundaries.

Support:

AWS
 ├── account A
 └── account B

Azure
 ├── subscription A
 └── subscription B

GCP
 ├── project A
 └── project B

This would make cross-account/resource-discovery/IAM testing much more realistic.


36. Plugin Architecture

As CloudEmu grows, providers/services should ideally be extensible.

A future architecture could allow:

CloudEmu Core
     │
     ├── AWS plugin
     ├── Azure plugin
     ├── GCP plugin
     ├── Kubernetes plugin
     └── custom provider

This prevents the core from becoming a monolith.


37. MCP Integration

An MCP server could expose CloudEmu to AI agents.

For example:

AI Agent
   ↓
MCP
   ↓
CloudEmu
   ↓
AWS/Azure/GCP simulation

An AI agent could then:

create infrastructure
inspect resources
modify networking
simulate failures
check costs
query topology
run tests

without touching real cloud accounts.

This could become a very interesting future direction.


38. Security Boundary

CloudEmu should make it extremely difficult to accidentally interact with real cloud infrastructure.

Possible safety features:

--offline
--no-network
--simulation-only

and clear startup messages:

CloudEmu is running in simulation mode.
No real cloud resources will be created.

This is particularly important when developers use real cloud SDK credentials on their machines.


39. Developer Experience Goal

The final experience should ideally be:

brew install cloudemu

cloudemu init my-project

cd my-project

cloudemu start

cloudemu apply cloudemu.yaml

go test ./...

cloudemu dashboard

And everything should be local.

No real AWS/Azure/GCP resources.

No unexpected cloud bills.

No complicated setup.


40. Priority

Not every issue needs to be solved immediately.

A practical priority would be:

P0 — Foundation

Persistence
Snapshots
Lifecycle CLI
Profiles
Environment isolation
Capability manifest
Documentation consolidation

P1 — Developer Experience

Dashboard
Topology visualization
Resource explorer
Seed fixtures
Terraform/OpenTofu integration
Better CI/Testcontainers experience

P2 — Simulation Depth

IAM enforcement
Advanced Kubernetes simulation
Multi-region
Multi-account
Better error fidelity
Rate limiting
Advanced chaos

P3 — Advanced Platform

Time travel
Cross-cloud cost comparison
Recording/replay
MCP
Plugin architecture
AI-assisted simulation

Final Goal

CloudEmu should not try to become simply:

"another LocalStack."

Its stronger long-term identity is:

A deterministic, programmable, multi-cloud infrastructure simulation platform for AWS, Azure, GCP and Kubernetes.

The core already provides many of the difficult building blocks:

AWS ─────┐
Azure ───┤
GCP ─────┤
K8s ─────┤
         ↓
   CloudEmu Engine
         │
 ┌───────┼────────┬────────┐
 ↓       ↓        ↓        ↓
Network Cost    Chaos   Metrics
 ↓       ↓        ↓        ↓
Topology IAM   Failures Monitoring
         │
         ↓
    Test / Simulate

The next major step is turning this powerful engine into a persistent, reproducible, observable and extremely easy-to-use developer platform.

The objective should be:

From "powerful cloud emulator" → "complete local cloud laboratory."

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions