DE
- AI acceleration
- Industries
- Finance
Nearshore software development for finance—secure, scalable, and compliant solutions for banking, payments, and APIs.
- Retail
Retail software development services—e-commerce, POS, logistics, and AI-driven personalization from nearshore engineering teams.
- Manufacturing
Nearshore manufacturing software development—ERP systems, IoT platforms, and automation tools to optimize industrial operations.
- Finance
- What we do
- Services
- Software modernization services
- Cloud solutions
- AI – Artificial intelligence
- Idea validation & Product development services
- Digital solutions
- Integrations for digital ecosystems
- A11y – Accessibility
- QA – Test development
- Technologies
- Front-end
- Back-end
- DevOps & CI/CD
- Cloud
- Mobile
- Collaboration models
- Collaboration models
Explore collaboration models customized to your specific needs: Complete nearshoring teams, Local heroes from partners with the nearshoring team, or Mixed tech teams with partners.
- Way of work
Through close collaboration with your business, we create customized solutions aligned with your specific requirements, resulting in sustainable outcomes.
- Collaboration models
- Services
- About Us
- Who we are
We are a full-service nearshoring provider for digital software products, uniquely positioned as a high-quality partner with native-speaking local experts, perfectly aligned with your business needs.
- Meet our team
ProductDock’s experienced team proficient in modern technologies and tools, boasts 15 years of successful projects, collaborating with prominent companies.
- Why nearshoring
Elevate your business efficiently with our premium full-service software development services that blend nearshore and local expertise to support you throughout your digital product journey.
- Who we are
- Our work
- Career
- Life at ProductDock
We’re all about fostering teamwork, creativity, and empowerment within our team of over 120 incredibly talented experts in modern technologies.
- Open positions
Do you enjoy working on exciting projects and feel rewarded when those efforts are successful? If so, we’d like you to join our team.
- Hiring guide
How we choose our crew members? We think of you as a member of our crew. We are happy to share our process with you!
- Rookie boot camp internship
Start your IT journey with Rookie boot camp, our paid internship program where students and graduates build skills, gain confidence, and get real-world experience.
- Life at ProductDock
- Newsroom
- News
Stay engaged with our most recent updates and releases, ensuring you are always up-to-date with the latest developments in the dynamic world of ProductDock.
- Events
Expand your expertise through networking with like-minded individuals and engaging in knowledge-sharing sessions at our upcoming events.
- News
- Blog
- Get in touch
08. Oct 2026 •10 minutes read
Two systems, one number, and no rule that gets it right
Danijel Dragičević
Software Engineer
What does it take to trust an agent’s decision and then prove months later that it was reasonable?
The problem nobody has a clean rule for
Two systems hold a number for the same thing, and the numbers don’t match.
A warehouse management system reports 40 units of a product on the shelf. The e-commerce platform says 32 are available to sell. Both are functioning as intended and are authoritative within their own domains, and neither is lying.
The reflex is a reconciliation rule. Always trust the warehouse. Trust the newer timestamp. Flag anything above a 5% delta. I’ve written versions of all three, and they rot the same way:
A fixed rule encodes exactly one cause, but a discrepancy of the same size and shape can have many.
Let’s look at six completely ordinary reasons why two systems might drift apart:

Look at the stale-sync and in-transit-return cases: the warehouse is right in both. Now the reserved stock and quarantine cases: same systems, same gap shape, warehouse wrong both times. “Trust the newer timestamp” gets the first right and the second wrong. “Trust the warehouse” gets the quarantine wrong. Then there’s the unexplained case, where nothing accounts for the gap, and the frank answer is “I don’t know, a human needs to look at this.” A 5% threshold would flag it, but it would also flag the other ones, proving that a fixed rule can’t distinguish an explainable gap from one it can’t account for.
None of these limitations means that hardcoded rules are useless. Most conflicts really do have a single cause, and where they do, a rule that encodes it is cheaper and easier to reason about than anything I’m about to build. The real challenge lies in what’s left over: the cases where a fixed rule quietly fails, which ends up as somebody’s month-end reconciliation work.
The person who does this job well isn’t applying a rule at all. They open the movement history, check what shipped, look at when each system last updated, form a hypothesis, and, on a good day, say, “This one’s weird, I’m escalating it.” That’s judgment over a small, bounded set of facts, exactly the shape of a task an LLM with tools is good at. So I built one to find out what the obvious objections cost in code.
The architecture
Three small services, no shared database, no shared types package:

global-wms-service and retail-storefront-service are deliberately dumb Express apps holding in-memory data. They model two genuinely independent systems: separately owned, separately shaped, reachable only over HTTP. They are supposed to disagree.
resolver-agent exposes POST/resolve/:sku, hands the SKU to Claude with a small set of tools, and writes the outcome to a SQLite database.
1. How the agent actually decides
The agent gets one instruction: investigate this SKU, plus five read-only tools and one for recording its verdict. Nothing tells it which tools to call, in what order, or when it has enough to conclude. It works all of that out from the tool descriptions.
A description that only names the endpoint gives it nothing to reason with. “Get warehouse movements for a SKU” is accurate but useless. It states what comes back, not what the result means or when it would change the answer. The same tool, written to carry that:
// services/resolver-agent/src/agent/tools.ts
const getWmsMovements = betaZodTool({
name: "get_wms_movements",
description: `Get the warehouse's movement history for a SKU (cycle counts, manual adjustments,
quarantine events, and quarantine releases), most recent first. Use this to find events
that happened after the last cycle count and that would change the real sellable
quantity — most importantly quarantine events, which reduce sellable stock immediately
even though the on-hand count only catches up at the next cycle count.`,
inputSchema: z.object({
sku: z.string().describe("The SKU to look up, e.g. SKU-001."),
}),
run: async ({ sku }) => JSON.stringify(await fetchJson(`${WMS_URL}/movements/${sku}`)),
});
Instead of just documenting the endpoint, this description provides the exact domain knowledge needed to keep the agent from blindly guessing. Each of the five reading tools carries a sentence like this:
- get_wms_stock: “This count has NOT been adjusted for anything found in get_wms_movements since then.“
- get_storefront_stock: “already has pending-order reservations subtracted and in-transit return credits added — it is not a physical count.“
- get_pending_orders: “Units reserved this way may still be physically present at the warehouse until they ship.“
- get_returns_in_transit: “Until a return like this lands, expect the warehouse and webshop numbers to disagree by exactly its quantity.“
The system prompt is much shorter, and it does two jobs. It lists the four explanations worth checking, and it tells the agent that finding none of them is a valid answer rather than a failed run:
/// services/resolver-agent/src/agent/runner.ts
const SYSTEM_PROMPT = `You are a stock discrepancy resolver. ...
Investigate using your tools before concluding anything — check for sync lag (compare
timestamps), pending orders, returns in transit, and quarantine events.
If nothing you find explains the gap, don't guess — set recommendedAction to
needs_human_review and say plainly what you checked and why nothing explained it.
...
Always finish by calling record_resolution exactly once.`;
const runner = client.beta.messages.toolRunner({
model: "claude-sonnet-5",
max_tokens: 4096,
thinking: { type: "adaptive" },
output_config: { effort: "medium" },
system: [{ type: "text", text: SYSTEM_PROMPT, cache_control: { type: "ephemeral" } }],
tools,
cache_control: { type: "ephemeral" },
max_iterations: 10,
messages: [{ role: "user", content: `Investigate SKU ${sku}. ...` }],
});
The SDK’s tool runner owns the agentic loop: call the model, run whichever tools it asks for, feed the results back, repeat. max_iterations: 10 is the hard stop if something goes sideways.
{
"sku": "SKU-004",
"wmsQty": 40,
"storefrontQty": 32,
"resolvedQty": 32,
"confidence": "high",
"recommendedAction": "trust_storefront",
"reasoning": "Warehouse physical count was 40 as of the 07:00:31 cycle count. A quarantine
event at 12:00:31 pulled 8 damaged units aside (found during spot inspection), which
reduces sellable stock immediately but hasn't yet been applied to quantityOnHand (that
only happens at the next cycle count). 40 - 8 = 32, exactly matching the webshop's current
available quantity of 32. No pending orders or in-transit returns exist to further adjust
the number. The gap is fully explained by the quarantine lag, so the webshop's figure of
32 is the correct current sellable quantity."
}
Notice the phrase right before the conclusion: “No pending orders or in-transit returns exist to further adjust the number.” It found its explanation early and still checked the two hypotheses that would have contradicted it. I can verify that, because the downstream services log every request they receive:
[global-wms-service] GET /inventory/SKU-004 -> 200 (1ms)
[global-wms-service] GET /movements/SKU-004 -> 200 (11ms)
[retail-storefront-service] GET /stock/SKU-004 -> 200 (1ms)
[retail-storefront-service] GET /returns/in-transit?sku=SKU-004 -> 200 (0ms)
[retail-storefront-service] GET /orders/pending?sku=SKU-004 -> 200 (0ms)
The agent executed five read requests to check every possible explanation. And it did this for all scenarios, in varying order, never stopping early once it had a sufficient explanation. That costs about 15 seconds and five HTTP calls per SKU, and the payoff is verifiable trust. If the agent claims it found no underlying cause, you can check the logs to confirm it actually looked, rather than assuming it just hallucinated a thorough-sounding excuse. A ledger entry stating “no pending orders were found” means a request was actually sent and returned empty.
2. What the descriptions cost, and what prompt caching saves
Measured via the API, the system prompt and six tool descriptions cost 2,218 tokens. The tools alone consume 1,898 tokens, with record_resolution eating about 500 tokens.
This highlights the main catch: an agentic loop doesn’t just send its prompt once. The quarantine case required three separate requests, re-sending the system prompt and tools every time. That token cost is a rounding error for a six-SKU test on my laptop, but with a few thousand discrepancies across a real catalog, it can become a serious problem.
The fix is prompt caching. You place a marker in the prompt (cache_control), and the API keeps everything before that marker in its already-processed form for a few minutes. Any later request that starts with exactly the same content reads it back instead of processing it again, at roughly a tenth of the price. Everything hinges on the word starts. It’s a prefix match from the very first byte, so any changes between requests have to be appended at the end, after everything that stays the same.
// services/resolver-agent/src/agent/runner.ts
system: [{ type: "text", text: SYSTEM_PROMPT, cache_control: { type: "ephemeral" } }],
tools,
cache_control: { type: "ephemeral" },
messages: [{ role: "user", content: `Investigate SKU ${sku}. ...` }],
The first marker sits on the system prompt and covers the tool definitions as well. The API always assembles the prompt as tools, then system, then messages, so a marker on the system block catches both. That’s 2,218 tokens identical in every request this service will ever make. The SKU being investigated lives in the user message, which comes last. A marker at the end would store a chunk ending with “investigate SKU-004,” and the next request would ask about SKU-005. That doesn’t match, so every SKU would store its own copy, and none of them would ever be read back.
The second marker is the standalone cache_control, which moves with the conversation as the tool’s results pile up. That’s what makes the second and third requests for a single resolution cheap: each one reads back the previous turns, rather than paying for them again.
Caching fails silently. If anything disturbs the prefix (a timestamp added to the system prompt, a tool built differently on one request), it simply stops matching, and you quietly go back to paying full price with no error to tell you. The only way to know if it’s still working is to check the numbers the API returns with every response. So, I added the log statement for each resolution:
SKU-001 -> trust_wms [cache: 5174 read, 3389 written]
SKU-002 -> trust_storefront [cache: 7431 read, 1422 written]
SKU-003 -> trust_wms [cache: 7392 read, 1306 written]
SKU-004 -> trust_storefront [cache: 7408 read, 1445 written]
SKU-005 -> needs_human_review [cache: 7320 read, 1244 written]
SKU-006 -> needs_human_review [cache: 7218 read, 1053 written]
The first resolution pays to store the whole prefix (3,389 tokens written), and every one after it reads that prefix back and stores only its own conversation (~1,300). Across the run, that’s roughly two-thirds less input costs.
3. What the agent is not allowed to do
Both mock services expose full POST, PUT, and DELETE endpoints, and I use them to build test scenarios by hand. The agent has never called one, and cannot, for the obvious reason:
// services/resolver-agent/src/agent/tools.ts
return [
getWmsStock,
getWmsMovements,
getStorefrontStock,
getPendingOrders,
getReturnsInTransit,
recordResolution
];
There is no updateWmsStock tool or some generic httpRequest tool. The boundary isn’t a permission check that could be misconfigured or a server-side rule with a gap – it’s the absence of a capability. Each read tool is a hardcoded GET against a specific path; none takes a URL or a method as a parameter. The access log confirms this from the other side: every request the agent has ever generated was a GET, and the only writes were those I issued by hand with curl. The only tool that has side effects is record_resolution; it takes the resulting decision and stores it in a database.
// services/resolver-agent/src/agent/tools.ts
const recordResolution = betaZodTool({
name: "record_resolution",
description: `Record how you resolved this SKU. Call this exactly once, after you've investigated,
to finish the task. This both reports your decision and writes it to the permanent
audit ledger — there is no other way to submit a result.`,
inputSchema: z.object({
sku: z.string(),
resolvedQty: z.number().describe("The quantity you believe is actually correct and sellable right now."),
confidence: z.enum(["high", "medium", "low"]),
recommendedAction: z
.enum(["trust_wms", "trust_storefront", "needs_human_review"])
.describe("... needs_human_review: nothing you found explains the gap — don't guess, say so."),
reasoning: z.string().describe(
`Plain-language explanation of what you checked and why you reached this conclusion.
This is what makes the ledger entry auditable — write it for someone
who wasn't watching.`,
),
}),
run: async (input) => {
/* ... */
},
});
The tool’s design relies on three specific choices:
- Reporting and persisting are the same action. No separate “save” step the model could forget, and no path where it announces a conclusion that never reaches the ledger.
- needs_human_review is an enum value, sitting alongside the two real answers rather than being an error path.
- The facts in the ledger are not supplied by the model. When record_resolution runs, it re-fetches both systems itself:
// services/resolver-agent/src/agent/tools.ts
const [warehouse, webshop] = await Promise.all([
fetchJson(`${WMS_URL}/inventory/${input.sku}`, { treat404AsNull: true }),
fetchJson(`${STOREFRONT_URL}/stock/${input.sku}`, { treat404AsNull: true }),
]);
insertLedgerEntry({
wmsQty: warehouse?.quantityOnHand ?? null,
storefrontQty: webshop?.quantityAvailable ?? null,
resolvedQty: input.resolvedQty,
confidence: input.confidence,
recommendedAction: input.recommendedAction,
reasoning: input.reasoning,
...
});
The model supplies the judgment columns; the tool supplies the fact columns. If the model hallucinated “the warehouse reported 90 units“, the ledger would still record the 40 the warehouse actually returned, and the contradiction would sit visible in the same row.
4. The case that justifies the whole design
The above-mentioned scenarios are the starting point. The real test is the unexplained case. Imagine the warehouse shows 25 units, the webshop shows 15, and there is no clear reason for the difference. The timestamps are identical, there are no pending orders or returns, and the logs only show a single routine cycle count:
recommendedAction: needs_human_review
confidence: low
reasoning: Warehouse shows 25 on hand (last counted 2026-08-08T12:00:31Z), webshop
shows 15 available (synced at the exact same timestamp) — a 10-unit gap.
Checked warehouse movements: only event is the routine cycle count itself
(-1 delta, noted as normal variance), no quarantine events. Checked pending
orders: none (so no reservations explain webshop being lower). Checked
returns in transit: none (so no in-transit credit explains webshop being
lower, and that would push webshop higher, not lower, anyway). None of the
usual explanations (sync lag, pending reservations, in-transit returns,
quarantine) account for the 10-unit discrepancy. Since nothing found
explains the gap, this needs human review rather than a guess.
It didn’t just confirm the list was empty; it also realized that even if a return existed, it would skew the numbers the wrong way. That’s the difference between blindly checking a list and actually reasoning about a hypothesis.
The missing-record case asks for the same restraint as the agent. The product exists on the webshop with 10 units, and the warehouse has never heard of it. Both warehouse tools return 404 Unknown SKU, which surfaces to the model as a real error rather than an empty result:
[resolver-agent] investigating SKU-006
[resolver-agent] Unknown SKU: SKU-006
[resolver-agent] Unknown SKU: SKU-006
[resolver-agent] SKU-006 -> needs_human_review (confidence: low)
The tempting failure is subtle: there’s only one number available, so why not use it?
This isn't a simple sync-lag or quarantine discrepancy — the warehouse system of record has no data for this product whatsoever, so there's no physical count to reconcile against. [...] I'm not treating the webshop's 10 as a verified physical count.
A missing record is an entirely different class of error that the ledger structurally flags by setting wmsQty to null instead of zero. This refusal to guess is what makes the ledger valuable. A system forced to always output a number gives you no way to distinguish a real investigation from an invention.
That data is also worth tracking over time. I haven’t built this, but the obvious next feature is a simple tally of explained versus unexplained monthly discrepancies. A spike in needs_human_review means an upstream system is drifting; a drop shows that your integration work is succeeding.
5. Testing a system that never returns the same string twice
This is where agentic features stop resembling normal software and where directly asserting on model output often leads to flaky tests. What worked was drawing a hard line between the deterministic parts and the judgment.
Everything that isn’t the model is normal software. The tools are just functions that can be mocked and tested like any HTTP client:
// services/resolver-agent/src/agent/tools.test.ts
it("get_wms_stock throws the service's real error message on a 404", async () => {
mockFetchResolvedOnce(404, { error: "Unknown SKU: SKU-999" });
const getWmsStock = getTool(
buildTools("claude-sonnet-5", () => {}),
"get_wms_stock",
);
await expect(getWmsStock.run({ sku: "SKU-999" })).rejects.toThrow("Unknown SKU: SKU-999");
});
That test encodes a design decision: a failing tool must hand the model the service’s own error message, not a generic one. “Unknown SKU: SKU-999” is something the agent can reason about (and did, on the missing-record case). “Error” is not.
The HTTP layer is tested with resolveSku mocked out, and both mock services have contract tests that check their real responses against what their openapi.yaml promises. All of it runs in CI on every push, with no API key and no real calls to the Claude API:
resolver-agent 13 passed (13)
global-wms-service 15 passed (15)
retail-storefront-service 24 passed (24)
The model’s judgment gets an eval instead, which deliberately does not run in CI. It costs real tokens, and it asserts on the decision rather than the prose:
# scripts/eval.sh — checks recommendedAction against the documented expectation
expected_action() {
case "$1" in
SKU-001|SKU-003) echo "trust_wms" ;;
SKU-002|SKU-004) echo "trust_storefront" ;;
SKU-005|SKU-006) echo "needs_human_review" ;;
esac
}
PASS SKU-001: trust_wms
PASS SKU-002: trust_storefront
PASS SKU-003: trust_wms
PASS SKU-004: trust_storefront
PASS SKU-005: needs_human_review
PASS SKU-006: needs_human_review
All 6 scenarios resolved as expected.
./scripts/eval.sh 1:30.95 total
Six scenarios, 91 seconds. I run it by hand after touching the system prompt, the tool descriptions, or the seed data. That distinction to assert on the decision and not the wording is the whole trick. The prose is non-deterministic. The decision, with well-described tools, is far more stable than the “it’s an untestable black box” complaint suggests.
Conclusion
In the end, this project built an investigative process that creates a verifiable record rather than just acting as a smarter tiebreaker. The agent knows no more about warehouses than a standard rule does. It simply knows which questions to ask, asks them every time, and writes down its findings so a human can verify them later. That’s why the resolvedQty column matters less than the reasoning column. Six months after a bad number surfaces, auditors want to know why the decision was made and whether it was reasonable, rather than just asking what the system decided.
The full project is on GitHub. Clone it, start the three services, and point it at SKU-005 (the one where the interesting behavior is a refusal). Then build a scenario of your own and see whether it holds up. That’s a ten-minute experiment and considerably more convincing than my six.
If you’re weighing where agents genuinely fit in your own systems and where a plain rule is still the better engineering call, we’d be glad to help you think it through. Get in touch to discuss your specific use case.
Struggling with data consistency across your enterprise platforms?
Resolving complex edge cases in distributed architectures requires deep technical expertise. Explore our integration services to see how we build resilient, tightly synchronized, and scalable digital ecosystems that keep your data accurate.
Tags:Skip tags
Danijel Dragičević
Software EngineerDanijel Dragičević is a software developer and content creator who has been part of our family since April 2014. With a strong background in backend development, he has spent the past few years specializing in building robust services for API integrations. Passionate about clean code and efficient workflows, he continuously explores new technologies to enhance development processes.