What Makes an API Scalable? A Solution Architect's Perspective

"Our API needs to scale." I hear this sentence in almost every kick-off meeting. It usually means one of three very different things: the API should survive ten times more traffic, the API should survive ten times more consumers and teams, or the API should survive the next five years without a rewrite. Most systems I am asked to rescue were built for exactly one of these - and failed on the other two.

Since 2012 I have designed, extended and stabilized APIs for e-commerce platforms, enterprise systems, SaaS products and public-sector applications, mostly in PHP and Symfony. This article is the mental model I use as a solution architect when I review an API or design a new one. It is not a list of HTTP status codes. It is about the decisions that determine whether an API grows with your business or becomes the thing that slows it down.

A scalable API is one where growth - in traffic, in consumers, in features, in team size - costs you linearly or less. Everything else is just an API that has not been tested by success yet.

The three dimensions of scalability

Before talking about caches and queues, it helps to separate the problem. When I assess an API, I score it on three axes:

Dimension The question What failure looks like
Load Can it handle 10x the requests by adding hardware? Timeouts on sale days, a database at 100% CPU, "just buy a bigger server".
Change Can it add features and consumers without breaking existing ones? Every release needs coordination with five client teams. Nobody dares to rename a field.
Cost Does running and maintaining it get cheaper per request as it grows? Infrastructure bills grow faster than revenue. Onboarding a developer takes weeks.

Engineers tend to focus on load because it is measurable and fun to solve. Stakeholders feel change and cost, because that is where budgets disappear. A good architecture answers all three, and the order in which you address them is a business decision, not a technical one.

1. Scalability starts with the contract

An API is a promise. Every field you return, every status code you emit and every URL you publish becomes something another team or customer builds on. The moment you have consumers, your design mistakes are no longer yours alone - they are multiplied by the number of integrations.

This is why I start every API project with the contract, not with the database schema. Concretely:

  • Model resources, not procedures. POST /orders/{id}/cancellation scales, POST /cancelOrder does not. Resource-oriented APIs are predictable: once a consumer understands one resource, they understand all of them.
  • Write the OpenAPI specification first. It forces the conversation about naming, required fields and error formats to happen before code exists, when changes are free. With API Platform or a contract-first setup in Symfony, the specification and the implementation stay in sync automatically.
  • Standardize errors once. Use RFC 9457 Problem Details (application/problem+json) for every error, with a stable machine-readable type, a field-level breakdown for validation errors and a trace ID. Consumers write their error handling once instead of once per endpoint.
  • Be boringly consistent. One casing convention, one date format (ISO 8601, UTC), one pagination style, one way to express money (integer minor units plus currency). Inconsistency is not a style issue - it is a support cost that grows with every consumer.

None of this makes a single request faster. All of it decides how expensive your API becomes once twenty clients depend on it.

2. Stateless or stuck

The single most important property for horizontal scaling is statelessness: any request can be handled by any instance. If that holds, scaling for load is mostly an infrastructure problem - add instances behind the load balancer. If it does not hold, you are stuck scaling vertically, and vertical scaling has a ceiling and a price tag.

PHP actually gives you a head start here. Its shared-nothing execution model means every request starts clean. Yet I regularly find state hidden in places that quietly break horizontal scaling:

  • Server-side sessions on the local filesystem, so a user is "logged out" when the load balancer routes them to another node.
  • Uploaded files written to the local disk of the instance that received them.
  • In-memory or APCu caches used as a source of truth instead of a pure optimization.
  • Cron jobs that assume they run on exactly one machine.
  • Rate limiters and locks stored per instance, so the "limit" multiplies with every node you add.

The fix is always the same: authentication travels with the request (signed tokens or API keys), files go to object storage, shared state goes to a shared store such as Redis or the database, and scheduled work goes through a queue with proper locking. Once an instance can be killed at any time without anybody noticing, you are ready for autoscaling.

3. The database is where APIs actually fall over

In my experience, almost no API is slow because of PHP. It is slow because of what it asks the database. The application layer scales horizontally with little effort; the primary database does not. That makes data access the most important part of API performance work.

Kill N+1 queries before they kill you

An endpoint that returns 50 orders and loads the customer for each one separately sends 51 queries. In development with 10 rows, nobody notices. In production on a busy day, it is the endpoint that takes the whole platform down. Every list endpoint should have a fixed, small number of queries regardless of page size. I enforce this with tests that assert the query count of critical endpoints.

Offset pagination does not scale

LIMIT 20 OFFSET 200000 forces the database to read and discard 200,000 rows. It gets slower the deeper consumers page, and it produces duplicates or gaps when data changes between requests. For any collection that can grow without bound, I use keyset (cursor) pagination:

public function findPage(int $customerId, ?OrderCursor $cursor, int $limit): array
{
    $qb = $this->connection->createQueryBuilder()
        ->select('id', 'status', 'total_amount', 'currency', 'created_at')
        ->from('orders')
        ->where('customer_id = :customerId')
        ->setParameter('customerId', $customerId)
        ->orderBy('created_at', 'DESC')
        ->addOrderBy('id', 'DESC')
        ->setMaxResults($limit + 1); // one extra row tells us if there is a next page

    if (null !== $cursor) {
        $qb->andWhere('created_at < :createdAt OR (created_at = :createdAt AND id < :id)')
            ->setParameter('createdAt', $cursor->createdAt)
            ->setParameter('id', $cursor->id);
    }

    return $qb->executeQuery()->fetchAllAssociative();
}

Backed by a composite index on (customer_id, created_at, id), this query costs the same on page 1 and on page 10,000. The cursor is returned to the client as an opaque, encoded string, so you can change the implementation later without breaking anyone.

Separate reads from writes when the shapes diverge

Most APIs are read-heavy, often by a factor of 10 to 100. When the data a consumer wants to read looks very different from the data you write, do not force both through the same tables and joins. Read replicas, denormalized read tables or a search index (OpenSearch, Meilisearch) take load off the primary database and keep write paths simple. This is not CQRS as a religion - it is a pragmatic answer to a measured bottleneck.

4. Cache in layers, invalidate on purpose

The fastest request is the one that never reaches your application. I think about caching as a set of layers, from the outside in:

  1. The client - via Cache-Control and conditional requests.
  2. A CDN or reverse proxy - for public, shared responses.
  3. The application cache - for expensive computations and aggregated data.
  4. The database - its own buffer pool, which works best when the layers above take the repetitive load.

HTTP caching is criminally underused in APIs. Symfony makes conditional requests easy, and the important detail is to check freshness before doing the expensive work:

#[Route('/products/{id}', methods: ['GET'])]
public function show(int $id, Request $request): Response
{
    $lastModified = $this->products->lastModified($id) ?? throw $this->createNotFoundException();

    $response = new JsonResponse();
    $response->setLastModified($lastModified);
    $response->setEtag(hash('xxh128', $id . $lastModified->format(DATE_ATOM)));
    $response->setPublic();
    $response->setMaxAge(60);
    $response->setSharedMaxAge(300);

    if ($response->isNotModified($request)) {
        return $response; // 304, no serialization, no heavy queries
    }

    return $response->setData($this->productView->build($id));
}

The hard part of caching is never storing data - it is deciding when it is no longer true. My rule: every cache entry needs a documented owner and an invalidation strategy (time-based, event-based or tag-based) before it goes live. A cache without an invalidation plan is a bug report scheduled for later.

5. Do less inside the request

A scalable API keeps its synchronous path short. Anything that is slow, depends on a third party or can fail independently - generating PDFs, sending emails, calling a payment provider's reporting API, importing CSV files, running an LLM prompt - does not belong in the request-response cycle.

The pattern is simple: accept the work, persist the intent, hand it to a queue, and answer immediately with 202 Accepted and a resource the client can poll or subscribe to.

#[Route('/reports', methods: ['POST'])]
public function create(CreateReportRequest $input): JsonResponse
{
    $reportId = Uuid::v7();
    $this->reports->createPending($reportId, $input);
    $this->bus->dispatch(new GenerateReport($reportId));

    return new JsonResponse(
        ['id' => (string) $reportId, 'status' => 'pending'],
        Response::HTTP_ACCEPTED,
        ['Location' => $this->generateUrl('report_show', ['id' => $reportId])],
    );
}

With Symfony Messenger behind it, workers scale independently of the web tier. A traffic spike fills the queue instead of exhausting PHP-FPM workers, failed jobs are retried with backoff, and permanently failed ones land in a failure transport for inspection instead of disappearing. The API stays responsive even when the work behind it is not.

6. Design for retries, because they will happen

At scale, networks fail constantly. Mobile clients lose connectivity, load balancers time out, queues deliver messages more than once. A client that did not receive a response cannot know whether the operation happened, so it retries. If your API is not built for that, retries turn into duplicate orders and double charges.

GET, PUT and DELETE are idempotent by definition. For POST operations with side effects, I require an Idempotency-Key header. The important implementation detail: claim the key atomically in the database, not with a "check, then write" against a cache. Two concurrent retries must not both pass the check.

CREATE TABLE idempotency_keys (
    client_id     INT NOT NULL,
    idem_key      VARCHAR(64) NOT NULL,
    request_hash  CHAR(64) NOT NULL,
    status_code   SMALLINT NULL,
    response_body JSON NULL,
    created_at    DATETIME NOT NULL,
    PRIMARY KEY (client_id, idem_key)
);
public function claim(int $clientId, string $key, string $requestHash): ?StoredResponse
{
    try {
        $this->connection->insert('idempotency_keys', [
            'client_id'    => $clientId,
            'idem_key'     => $key,
            'request_hash' => $requestHash,
            'created_at'   => date('Y-m-d H:i:s'),
        ]);

        return null; // first time we see this key: process the request
    } catch (UniqueConstraintViolationException) {
        $row = $this->connection->fetchAssociative(
            'SELECT request_hash, status_code, response_body FROM idempotency_keys WHERE client_id = ? AND idem_key = ?',
            [$clientId, $key],
        );

        if ($row['request_hash'] !== $requestHash) {
            throw new IdempotencyKeyReused();      // same key, different payload: 422
        }

        if (null === $row['status_code']) {
            throw new IdempotentRequestInProgress(); // original still running: 409
        }

        return StoredResponse::fromRow($row);      // replay the original response
    }
}

The primary key does the concurrency control for you. After processing, the response is stored against the key, and a scheduled job purges keys older than the retention window you documented. The same thinking applies to consumers of your webhooks and queue messages: every handler must be safe to run twice.

7. Protect the system from its own success

An API that is open to the world will eventually receive more traffic than it can handle - from a marketing campaign, a misbehaving integration or someone who wrote a loop without a sleep. Scalable systems degrade gracefully instead of collapsing. Three mechanisms matter most.

Rate limiting per consumer

Limits protect the platform and make capacity plannable. They also turn into a product feature: tiers with different limits are a classic monetization lever. Symfony's RateLimiter component provides a token bucket out of the box:

# config/packages/rate_limiter.yaml
framework:
    rate_limiter:
        api_per_client:
            policy: 'token_bucket'
            limit: 100
            rate: { interval: '1 second', amount: 10 }
$limit = $this->apiPerClientLimiter->create($clientId)->consume();

if (!$limit->isAccepted()) {
    throw new TooManyRequestsHttpException($limit->getRetryAfter()->getTimestamp() - time());
}

Two details separate a demo from production: the limiter storage must be shared across all instances (Redis, not the local filesystem), and every response should expose the remaining budget in headers so well-behaved clients can throttle themselves before they hit 429.

Timeouts and circuit breakers for every dependency

Your API is only as available as its slowest dependency. A payment provider that responds in 30 seconds instead of 300 milliseconds will hold your PHP-FPM workers hostage until the whole API stops answering - including endpoints that never talk to that provider. Every outbound call needs an explicit, short timeout (timeout and max_duration in Symfony HttpClient), and dependencies that fail repeatedly should be short-circuited for a while instead of being hammered with more requests.

Bulkheads

Isolate workloads so that one cannot starve the other. Separate worker pools for critical and non-critical queues, separate PHP-FPM pools for public and internal traffic, separate database connections for reporting. When the CSV export goes wild, checkout must keep working.

8. Evolve without breaking anyone

The change dimension of scalability is mostly about one discipline: never break consumers by surprise. In practice that means:

  • Additive changes are free. New endpoints, new optional fields and new optional parameters do not need a new version. Consumers must be built to ignore unknown fields - put that into your API guidelines from day one.
  • Breaking changes are expensive and must be rare. Removing or renaming fields, changing types, making parameters required. These need a new version.
  • Expand, then contract. Add the new field next to the old one, migrate consumers, mark the old field deprecated, and only then remove it. Use the Deprecation and Sunset headers so clients can detect it programmatically.
  • Version the representation, not the business logic. v1 and v2 should be different serializers on top of the same services, not two copies of the application that drift apart.
  • Know your consumers. Log which client uses which version and which deprecated fields. You cannot sunset what you cannot measure.

9. Security has to scale with the surface

Every new endpoint is new attack surface. The most common API vulnerability in the OWASP API Security Top 10 is Broken Object Level Authorization: an endpoint loads /orders/4711 without checking whether the caller is allowed to see order 4711. It is rarely a knowledge problem - it is a scaling problem. With 200 endpoints and ten developers, someone will eventually forget the check.

So the answer cannot be "be careful". It has to be structural:

  • Authorization decisions live in one place, for example Symfony Voters, and controllers declare them with #[IsGranted('ORDER_VIEW', 'order')] instead of hand-written if statements.
  • Multi-tenant data access is scoped in the repository layer, so a query without a tenant is impossible rather than forbidden.
  • Responses are built from explicit output models, never by serializing database rows or entities directly. New columns do not leak automatically.
  • Automated tests try to access other users' resources for every endpoint, not just the ones someone remembered.

Security that depends on individual discipline does not scale. Security that is the default path does.

10. You cannot scale what you cannot see

Every scaling decision should start with data. Before I recommend a cache, a queue or a new database, I want to know where the time actually goes. The minimum I put into every production API:

  • Latency percentiles per endpoint - p50, p95 and p99. Averages hide exactly the slow requests your customers complain about.
  • Error rates split by 4xx and 5xx - a rising 4xx rate often means a consumer broke, a rising 5xx rate means you did.
  • A correlation ID on every request, propagated to logs, queue messages and outbound calls, and returned in the response so support can find the exact request a customer is talking about.
  • Structured logs in JSON, not free text, so they can be queried instead of grepped.
  • Saturation metrics - queue depth, worker utilization, database connections. These warn you before latency does.

Observability is also what makes an architecture defensible in front of stakeholders. "We need Redis" is an opinion. "62% of p99 latency on the checkout endpoint is spent recomputing the same shipping rates" is a decision.

The part that is not about code

Everything above is engineering. What turns it into architecture is knowing which of it you need now, which you need later and which you will never need. The most expensive APIs I have seen were not under-engineered - they were over-engineered for a scale that never came, while the actual bottleneck sat untouched.

As a solution architect, these are the principles I hold myself to:

  • Start with a modular monolith. Microservices solve organizational scaling problems, and they introduce distributed-systems problems in exchange. Split a service out when a team boundary or a measured load profile demands it, not because it looks good on a diagram.
  • Prefer boring technology. PHP, Symfony, MySQL or PostgreSQL, Redis and a message queue carry an enormous amount of traffic when used well. Every additional technology is an operational cost your team pays forever.
  • Make decisions explicit. Architecture Decision Records document why something was built the way it was. They make onboarding faster and prevent the same discussion from happening every six months.
  • Translate between business and engineering. A deadline, a budget and a growth forecast are architectural inputs. The right design for a startup validating a market is wrong for a platform processing millions of orders, and vice versa.

A checklist for decision makers

If you are responsible for a product that depends on an API, these questions tell you quickly where you stand. You do not need to understand the implementation to ask them - but you should be worried if nobody on the team can answer them.

Question Why it matters
Can we add a server and get more capacity, without code changes? If not, growth means a project instead of a configuration change.
Is there a written, versioned API specification? Without it, every integration depends on tribal knowledge.
What happens if a client sends the same payment request twice? Duplicate charges are a direct financial and reputational risk.
What happens when our most important third-party provider is down? The answer should be "that feature degrades", not "everything stops".
What is our p99 latency on the three most important endpoints? If nobody knows, performance problems are discovered by customers.
How do we change a field without breaking our partners? The answer reveals whether the API can evolve or is frozen.
How is it guaranteed that a customer can never read another customer's data? "Developers check it" is not a guarantee.

Conclusion

A scalable API is not the result of one clever technology choice. It is the sum of many unspectacular decisions made consistently: a clear contract, stateless services, honest data access, layered caching, asynchronous work, safe retries, graceful degradation, careful evolution, structural security and real observability. Each of them is cheap when designed in early and expensive when retrofitted under pressure.

The good news is that most existing APIs do not need a rewrite to get there. They need someone who can find the two or three decisions that actually limit growth, fix those first and leave the rest alone.


API architecture for growing products

Is your API ready for the next 10x?

I help companies design new APIs and stabilize existing ones - from the first OpenAPI draft to production systems under real load. As a solution architect and hands-on senior engineer, I do not stop at the diagram: I work in your codebase with your team until the result runs in production.

API Architecture Review

A written assessment of your API across load, change and cost: bottlenecks, risks and a prioritized roadmap. Fixed scope, delivered in writing, understandable for engineering and management.

API Design and Implementation

Contract-first design and implementation of new APIs with PHP, Symfony and API Platform - including authentication, versioning, documentation and observability from day one.

Performance and Stabilization

Hands-on work on systems that are already struggling: query optimization, caching, asynchronous processing, rate limiting and resilience patterns - measured before and after.

hello@dogan-ucar.de View all engagements →

A response usually comes the same day. Remote and on-site, Germany and Europe.