All articles

Testing Microservices at Scale: Payloads and Contracts

One service is the easy part. API Testing: A Practical Guide covers testing a single API well: the test pyramid, schema validation, the unhappy paths, and combining API and UI checks in one test. This is the other half.

Real systems are dozens of services passing large messages to each other, caching for speed, aggregating data into tables other services read, and quietly changing shape under you. The bugs that cost the most money live in that machinery, not in a button.

Everything below is how I test it, from a year on a large retailer’s promotions and reviews services. The client stays anonymous; the techniques do not.

A note on the examples. The bugs, payloads, and workflows here are real, from production work on a large retailer’s promotions and reviews services. To show them safely I rebuild everything on Vendly, a fictional marketplace I use across this blog, so the structure stays deep and real without exposing any client’s data. The techniques are exactly what I did.

A real microservice request is not three fields

Part one uses small, friendly payloads, because that is how you learn. Real backend services are not friendly.

The first time I opened the request going into a promotions service on a retail checkout, the body was the entire bag: every line item with its price, quantity, category, and eligibility flags, the customer context, the coupon codes, and a denormalized block of product data per item.

One realistic “apply the promotions to this cart” call ran past two hundred fields, nested nine levels deep, over eighty kilobytes of JSON.

You cannot hand-type that, and should not: every value you invent might not be valid, and an invalid payload tests nothing except your patience.

Pull a real request body from the logs, then change one thing

When a payload is too big to author by hand, I do not author it. I find one that already happened.

Every serious backend writes its traffic to centralized logs, and every team has a tool for reading them. On my platforms that has been Google Cloud’s logging and monitoring tools, plus Grafana and Datadog; other teams live in Splunk or the ELK stack (Elasticsearch and Kibana). They all do the same job: search production traffic and open a single real request, headers and body and all.

With read access to an environment’s logs, you can find a request the service already received and copy the body straight out of it. Now you hold a payload the service has genuinely accepted, every field populated and valid.

From there, testing becomes surgical. I drop that real body into Postman, a Pytest fixture, or a Playwright request, and change only the handful of fields the case is about. Everything else stays valid.

Pull a real payload from the logs, then change only what the test is about. Everything else stays correct without you touching it.

Build the payload from a real request, not from a guess.Find a real request in the logsGoogle Cloud, Grafana, Datadog, Splunkneeds log read accessCopy the request bodyvalid, fully populated, exactly what the service acceptsChange only the fields under testeverything else stays valid for freeSend and assertagainst a fresh, realistic result

This is why log access is worth asking for on day one. Not just to debug a failure after the fact, but as the fastest way to build realistic test data for inputs too complex to invent.

I have written these suites in Java, in Python with Pytest, and in Postman’s JavaScript on the retail work. Java and Python are not my favorite, but it does not matter: the language changes, the thinking does not.

The examples here are in JavaScript because that is what I reach for now and what most QA teams can pick up and run today.

Do not hardcode the answer. Recompute it from the request.

The most common mistake I see in API tests for business logic: the test calls the endpoint, sees the discount came back as 132, and writes:

expect(body.discountTotal).toBe(132)

That assertion is brittle and, worse, dishonest.

  • Brittle: the moment someone changes the test data, it breaks for a reason that has nothing to do with a real bug.
  • Dishonest: if the service was wrong and actually owed the customer 138, you just copied the bug into your test and made it permanent.

A hardcoded expected value proves the response did not change. It does not prove the response is correct.

The fix is to compute the expected value yourself, from the inputs, and assert the service agrees. Now the test checks the arithmetic the service did, not a number someone once wrote down.

Here is the idea on a public sandbox you can run. DummyJSON serves carts where each line item carries a price, a quantity, and a discounted total, and the cart reports its own totals. Instead of trusting the cart’s numbers, recompute them from the items:

Playwright
import { test, expect } from '@playwright/test'

test('verify that the cart totals match the line items', async ({ request }) => {
  const response = await request.get('https://dummyjson.com/carts/1')
  const cart = await response.json()

  // Recompute the totals from the line items, do not trust the summary
  const expectedTotal = cart.products.reduce((sum, p) => sum + p.total, 0)
  const expectedDiscounted = cart.products.reduce((sum, p) => sum + p.discountedTotal, 0)

  expect(cart.total).toBeCloseTo(expectedTotal, 2)
  expect(cart.discountedTotal).toBeCloseTo(expectedDiscounted, 2)
})

Two details worth stealing:

  • I use toBeCloseTo rather than toBe because money is full of rounding, and a failure on the third decimal place is noise, not a finding.
  • I derive both expected numbers from the same data the service was given, so the test survives the data changing and still catches the service getting the math wrong.

On the real promotions work this pattern did the heavy lifting. The service ran tiered “spend more, save more” rules: pass one threshold and you unlock a bigger discount, pass the next and a bigger one again.

The test summed only the line items eligible for that promotion, worked out which tier that subtotal lands in, calculated how far the cart was from the next tier, then asserted the response matched all three.

When a tier boundary moved, or an item that should not have counted slipped into the total, the test caught it, because it checked the rule and not a remembered number.

This is not hypothetical. A tiered promotion came back 200 OK and a shallow status check was green, so everything passed at a glance.

But when the test recomputed the discount, the numbers did not hold up: the progress within the current reward band was impossible, a negative value where it must sit between zero and one, and three of the five recomputing assertions went red. The status was a 200, the shape was valid, and the math underneath was wrong.

A customer would have been charged the wrong amount. It only surfaced before release because the assertion checked the arithmetic against the rule instead of nodding at the 200.

POST▾{{baseUrl}}/offers/v2/baskets/bsk_7731/evaluateSend▾ParamsAuthorizationHeaders (10)BodyScriptsSettings● raw JSON ▾1{2”basketId”: “bsk_7731”,3”lines”: [4{5”sku”: “VND-APR-00417”,6”qty”: 2,7”unitPrice”: 18.008},9{10”sku”: “VND-HOM-02288”,11”qty”: 1,12”unitPrice”: 12.0013}14],15”couponCodes”: [“STACKUP”]16}BodyCookies (1)Headers (6)Test Results(2/5)200 OK377 ms1{2”countedSubtotal”: 48.00,3”appliedOffers”: [4{5”offerKind”: “SPEND_AND_SAVE”,6”offerCode”: “STACKUP”,7”rungReached”: 0,8”rungApplied”: true,9”withinRungProgress”: -51.97,10”toNextRung”: 52.00,11”savedAmount”: 0.0012}13]14}withinRungProgress must be 0 to 1. -51.97 is impossible.rungApplied is true, yet rungReached is 0 and savedAmount is 0.

Recreated on Vendly, my fictional marketplace. The offers call came back 200 OK, but Test Results read 2 of 5 and the ladder is impossible: withinRungProgress is negative and rungApplied is true while nothing was saved.

Set up the data, clear the cache, then assert

Fast services cache, and caches lie. A promotions service is not reading the database on every call, it is reading product and promotion data cached for speed. So if you load a new promotion and immediately fire your test, there is a good chance you are testing yesterday’s cache and calling it a pass.

The sequence for these tests was always three steps, in order:

  • Seed the data. Load the promotion and product setup the case needs.
  • Invalidate the cache. Call the service’s cache-reload endpoint so it forgets the old state.
  • Then make the call you actually want to test, and assert against the fresh result.

Skip the middle step and your test becomes a coin flip. Worse, the bugs that live here are the ones that hurt in production: a price updates but the old one keeps serving, an item is removed but still shows as eligible, the cache is cold after a deploy and the first customers get the wrong answer.

Those are not edge cases, they are Tuesday. Testing them on purpose means writing then reading through the cache, checking eviction, and checking what the service does when the cache is empty or unavailable.

This is the same trap I keep writing about from the UI side: the screen looks right because something downstream is serving a stale copy of the truth. One layer down it is just as real, and the cache is usually where it hides.

List endpoints hide their bugs in the query string

A read endpoint with filtering, sorting, pagination, and search looks like the easy part of the API. It is not. The happy path is trivial and the bugs all live in the parameters, exactly where most suites stop looking.

When I tested a reviews API, the happy-path GET was a handful of checks. The real suite was everything I could do wrong to the query string:

  • Pagination boundaries. An offset past the end of the data, a limit above the allowed maximum, and the values that should be rejected outright: zero, negative, and non-integer.
  • Sorting. Every key the API claims to sort by, in both directions, plus a few malformed sort tokens to confirm they are rejected and not silently ignored.
  • Filtering. Valid values, invalid values, and the same filter repeated twice to see which one wins.
  • Search. Empty terms, padded whitespace, mixed case, a term that matches nothing, and an injection probe like ' OR 1 to confirm the API treats it as literal text and never as something to run.

The rule that matters most on every one of those: assert the error shape, not just the status code. A good API does not only answer 400, it answers with a body that tells the caller what went wrong.

The standard is problem+json (RFC 9457, formerly RFC 7807), a small object with type, title, status, and a human-readable detail.

When the API rejects a bad limit, I want the detail to say something like “limit must be positive,” and I assert on exactly that, because a 400 with a useless body is its own bug.

I caught a sharp one right here. I sent an offer-evaluation request with a product id that did not exist, expecting a clean validation error. Back came a 500 whose body was an internal crash: a null-object method call, with the service’s own class and package names spilled into the response.

Two bugs in one. An unknown id should be a clean validation error, not a server crash, and the crash leaked the service’s internals to the caller, the kind of detail an attacker reads like a map.

Neither would have shown up if I had only checked the status code. That information-disclosure angle is its own subject, and I go deeper on it in security and authentication testing.

POST▾{{baseUrl}}/offers/v2/baskets/bsk_8842/evaluateSend▾ParamsAuthorizationHeaders (10)BodyScriptsSettings● raw JSON ▾1{2”basketId”: “bsk_8842”,3”lines”: [4{5”sku”: “VND-XXX-99999”,6”qty”: 17}8],9”couponCodes”: [“STACKUP”]10}BodyCookies (1)Headers (6)Test Results500 Internal Server Error41 ms1{2”timestamp”: “2026-03-09T19:24:50Z”,3”status”: 500,4”error”: “Internal Server Error”,5”path”: “/offers/v2/baskets/bsk_8842/evaluate”,6”message”: “Cannot invoke Item.unitPrice() because item is null”,7”trace”: “io.vendly.offers.engine.LadderEvaluator.scoreRung:212”8}An unknown sku should be a clean 422, not a 500 server crash.The stack trace leaks internal class paths to the caller.

Recreated on Vendly, my fictional marketplace. An unknown sku made the offers service throw and return a 500 that leaked an internal stack trace. It should have been a clean 422: a validation gap and an information leak in one response.

Here is the negative-testing shape against the DummyJSON sandbox, which is honest about one thing:

Playwright
import { test, expect } from '@playwright/test'

const api = 'https://dummyjson.com/products'

test('verify that a missing product returns a clean 404', async ({ request }) => {
  const response = await request.get(`${api}/0`)
  expect(response.status()).toBe(404)

  const body = await response.json()
  expect(body).toHaveProperty('message')
  expect(body.message).toContain('not found')
})

test('a negative limit should be rejected', async ({ request }) => {
  const response = await request.get(`${api}?limit=-5`)
  // A well-built API answers 400 here. This sandbox answers 200.
  // That gap is exactly the kind of permissive validation worth flagging.
  expect(response.status()).toBe(400)
})

Run that second test and it fails, because DummyJSON returns 200 for a negative limit instead of a 400. That is not a problem with the test, it is the test doing its job: it found that the API does not validate its pagination input.

On a real service I would file that, because permissive input handling on a list endpoint is how you end up serving the whole table to anyone who asks.

Write the assertion for the behavior you want, run it against what you have, and treat the gap as a finding rather than editing the test to match the bug.

Check the saved data, not just the reply

A passing response tells you what the service said, not what it stored. On the reviews and promotions work the data lived in more than one place at once: a primary store, a cache in front of it, and aggregate tables other services read from.

The API could answer correctly while the record behind it was wrong, or answer wrongly while the record was fine.

So for anything that writes or aggregates, I do not stop at the response. I read the data underneath and check they agree.

The shape is simple: act through the API, then read the source of truth directly and assert they match.

// 1. Act through the API
await api.post('/reviews/v1/items/90417/reviews', { stars: 5, title: 'Holds up' })

// 2. Read the source of truth directly, not the API's own echo
const stored = await db.collection('reviews').findOne({ itemId: 90417, title: 'Holds up' })
expect(stored.stars).toBe(5)

// 3. Check the summary the product page actually reads
const summary = await db.collection('reviewSummary').findOne({ itemId: 90417 })
expect(summary.averageStars).toBeCloseTo(expectedAverageAfterReview, 2)

The database client depends on your stack; the thinking does not. If a review is submitted, does the rating average the product page reads actually move, by the right amount? If a promotion is loaded, does the stored eligibility match what the evaluate call returns?

When the response and the record disagree, you have found the bug both the UI test and the happy-path API test walk straight past.

This is the same lesson as when the UI is not the real problem (coming soon), one layer lower: the screen can lie because the API is wrong, and the API can lie because the data underneath is wrong. Check the truth all the way down.

Where AI carries this, and where you still cannot leave the room

This layer is where AI genuinely helps. Hand it that eighty-kilobyte payload and the schema and it drafts the variations far faster than you would by hand: the missing field, the wrong type, the boundary on every numeric value.

It is good at contracts and integrations, good at combing through a large response and telling you what changed, good at spotting that a backend is not returning what it should.

Generating test data and cases against a real payload is its strong suit, and I go through that workflow in using AI to generate tests and test data.

What it cannot do is tell you which rules matter. It will not know that a discount dropping the cart below the free-gift threshold should remove the gift, because that is a business decision someone made in a meeting, not a pattern in the data. It will not know that a “passing” number is actually wrong.

That judgment, which rule to trust, which edge case is real, what is safe to ship, stays with you. I dig into that division in QA as the control layer for AI-assisted development.

Hand the AI the payload, the schema, and the data-combing. Keep the math, the risk, and the call on what is correct for yourself.

The real problem with microservices: nobody owns the contract

Everything so far makes one service trustworthy on its own. The last problem is between services, and no single team owns it.

Service A calls Service B. Team B ships a change to their response, all their own tests pass, and Team A’s service breaks in production because a field they relied on quietly disappeared.

End-to-end tests are supposed to catch this, but in a system with twenty services, spinning them all up together is slow, expensive, and flaky, and you still only exercise the combinations you happened to think of.

This is the problem contract testing solves, and it is the missing piece for most teams running microservices. Instead of testing services together, the consumer writes down exactly what it expects from the provider, that expectation becomes a contract, and the provider is tested against it in isolation.

No shared environment, no spinning up the whole world, and the break is caught before either side ships.

The broker is the single source of truth both teams write to and read from.Consumer testdefines expectationsConsumer teampublishes contract belowPact Brokerstores every contractconsumer publishes · provider readsProvider verifyreplays the contractProvider teamreads contract from brokercan-i-deployreads the broker to gate both deploys

Pact is the most established tool here, and libraries like Pactum cover the same idea. It is consumer-driven: the team consuming the API defines the interactions they actually rely on, and the provider has to prove it still satisfies them.

That direction matters. The contract describes real usage, not a wish list of everything the provider could theoretically return.

Here is the shape of a consumer test in Pact’s JavaScript library. The current top-level import is now PactV4; the PactV3 API shown here still ships and is a clean way to see the pattern.

const { PactV3, MatchersV3 } = require('@pact-foundation/pact')
const { like, eachLike } = MatchersV3

const provider = new PactV3({
  consumer: 'WebApp',
  provider: 'OrdersService',
})

describe('Orders API contract', () => {
  it('verify that the provider returns a list of orders for a user', () => {
    provider
      .given('user 42 has orders')
      .uponReceiving('a request for user 42 orders')
      .withRequest({ method: 'GET', path: '/users/42/orders' })
      .willRespondWith({
        status: 200,
        headers: { 'Content-Type': 'application/json' },
        body: eachLike({
          id: like('ord_123'),
          total: like(49.99),
          status: like('shipped'),
        }),
      })

    return provider.executeTest(async mockServer => {
      const orders = await getOrders(mockServer.url, 42)
      expect(orders[0]).toHaveProperty('status')
    })
  })
})

Look closely at the matchers. like(49.99) does not mean “the total must be 49.99.” It means “the total must be a number shaped like this.” The contract is about types and structure, not specific values, which is exactly what you want.

Running this test produces a contract file, published to a Pact Broker, and the provider team runs verification against it in their own pipeline. If they are about to break you, their build fails, not yours.

The piece that ties it together in CI is can-i-deploy. Before either service deploys, it asks the broker one question: given the contracts on record, is it safe for this version to go to production? If a consumer depends on something the new provider version no longer offers, the deploy is blocked.

That one gate removes the whole category of “Team B broke Team A and nobody knew until customers complained.” If you are wiring contract verification into your pipeline, it slots in alongside everything else in continuous testing in CI/CD (coming soon).

A few things I have learned the hard way with Pact:

  1. Keep provider states reproducible. The given('user 42 has orders') line maps to setup code on the provider side; if that setup is flaky, your contract verification is flaky too. Treat it like real test data setup.
  2. Contract test the interactions you actually use, not every endpoint. The value is in encoding real consumer behavior. A contract for an endpoint nobody calls is just maintenance you signed up for.
  3. Contract testing does not replace functional API tests. It confirms two services agree on the shape of their conversation. It does not check that the business logic inside an endpoint is correct. You still need the functional tests from the rest of this guide.

Where each kind of test belongs

Once you have all of these in your toolkit, the question becomes how often to run each one and where. Here is how the layers compare on what decides it: speed, what they catch, and how much it costs to keep them green.

What it gives youUnitAPI / functionalContractAPI + UI smoke
Runs in milliseconds✓✓✓✗
Survives a CSS or layout change✓✓✓~
Catches a broken response shape✗✓✓~
Catches one service breaking another✗✗✓~
Verifies the real user journey✗✗✗✓
Cheap to keep green✓✓~✗

✓ yes   ~ partially or with effort   ✗ no

Read across the rows and the testing strategy writes itself. Where each layer earns its place:

  • Unit and contract tests run on every commit, because they are fast and isolated.
  • Functional API tests run on every pull request, against a dedicated test environment.
  • A small set of combined API plus UI smoke tests covers the critical flows, run before deploys and on a schedule against staging.
  • Full end-to-end-across-services is reserved for a handful of genuinely critical journeys, because they are the expensive, slow tests, and you want as few as you can responsibly get away with.

Where to start

You do not need all of this at once. Pick the one endpoint that carries the most risk, usually the one that touches money or eligibility, and build its tests properly: a real payload from the logs, the expected result recomputed from the inputs, a cache-reload before you assert, and a check that the stored data agrees with the response.

One endpoint tested that way catches more than a dozen that only check for a 200.

Then, if you run microservices, add contract testing before the next time one team’s change quietly breaks another. Pick the two services that talk most, write one consumer contract, and wire can-i-deploy into both pipelines.

It is far less painful than the production incident it prevents, and once your team feels that first blocked deploy save them, they never want to ship without it.

To make it stick rather than stay a good intention, give it an owner and a finish line:

  • Owner: QA. Build the real-payload, recomputed-math, cache, and stored-data checks on the highest-risk endpoint. Done when: a wrong total, a stale cache, or a mismatched record fails the suite on every pull request.
  • Owner: Dev. Write the first consumer contract between the two chattiest services. Done when: the provider verifies against the published contract in its own pipeline.
  • Owner: EM. Wire can-i-deploy into both pipelines. Done when: a provider change that breaks a consumer contract blocks the deploy instead of reaching production.

New to API testing? Start with the fundamentals in API Testing: A Practical Guide, then come back here.

Want a sandbox to practise on? DummyJSON gives you products and carts to recompute and assert against, and the Pact workshops walk you through contract testing end to end.

Found it useful? Share it.
Julia Pottinger

Written by

Julia Pottinger

Hi, I'm Julia. I've been in QA for over a decade. I spend my days testing software and my own time building apps and games, and I write here to share what I learn, the practical, honest lessons you can actually use.

Comments 0

Share your thoughts, ask questions, or add to the conversation.

Be kind and constructive. Stay on topic. No spam or self-promotion.
Loading comments…