How to Investigate Payment Declines with Claude and Payneteasy MCP
A 9-step workflow for investigating payment decline spikes using Claude connected to the Payneteasy MCP server.
Meet us at conferences around the world
SBC Summit Lisbon
SiGMA Europe
A 9-step workflow for investigating payment decline spikes using Claude connected to the Payneteasy MCP server.
An approval rate falls by two percentage points overnight. The dashboard confirms it and stops there. The useful question is not whether AI can help. It is which questions to ask, in which order, and at which point the answer stops being a fact and becomes a guess.
What follows is the sequence a payment operations team can run with Claude connected to the Payneteasy MCP server. Nothing here retries a payment, changes a setting or moves money.
"Approvals are down" is not a metric. Before the first prompt, settle three things: which ratio you are measuring, over which population, and against which comparison window.
Use windows of equal length and matching weekday shape. A Monday-to-Sunday week compared with a five-day stretch will produce a movement that means nothing.
Approval rate for merchant <name>, project <name>, 1–7 September, compared with 25–31 August. Same currency, same endpoint. Return approved, declined and fraud-filtered counts, plus the ratio for each period.
Keep fraud-filtered separate from declined. Transactions stopped by the platform's fraud filters never reached an issuer, so they belong to a different owner and need a different fix. Merging the two buckets is the fastest way to send an investigation to the wrong team.
A scoped token defines which tools and data the assistant can reach, and the MCP server is read-only (how Payneteasy MCP secures AI access). Establish that boundary first. A dimension you cannot query does not announce itself. It quietly turns into an assumption three steps later.
List the Payneteasy MCP tools you can call and the filters each one accepts. Then tell me which of these I cannot answer with them: approval rate by issuer country; decline reason breakdown by processor; configuration change history; settlement and payout figures.
A rising decline count is ambiguous on its own. It moves when acceptance falls, and it also moves when traffic grows. Pull both sides of the ratio for both periods before drawing any conclusion.
For both periods, give me total attempts, approved, declined and fraud-filtered, with the grand total and the drill-down ids. Tell me whether the decline share moved, the attempt volume moved, or both.
This step often resolves the concern without further investigation. Volume grew, absolute declines grew with it, the ratio held, and nothing has broken.
If the ratio genuinely moved, the next task is to find where the movement is concentrated. Check the tool schema first (full capability list), because not every dimension works both as a filter and as a grouping. Work one dimension at a time. A single grouped answer is easier to check than a matrix.
Break the declines down by processor for both periods, then by issuer country, then by card type. For each value show both periods and the delta. Flag any segment whose decline share moved further than the overall average, and give me the attempt volume alongside it so I can ignore thin segments.
One caution: processors do not share a common vocabulary for statuses and decline reasons (why this matters in a multi-processor stack). Where the mappings are not comparable, compare each processor against its own history rather than against its neighbour.
Once the segment is narrow, ask for the reasons behind it rather than across the whole account.
For <processor>, <issuer country>, <card type> in the affected period, group declines by reason and show the same grouping for the comparison period. Which reasons account for the increase?
Reasons fall into three groups with different owners. Issuer-side responses — insufficient funds, card restrictions, "do not honour" — generally point at the cardholder or the issuing bank. Technical responses may point to the merchant integration, gateway, processor, acquirer, network connection or route and should be investigated separately from issuer declines. Fraud-filter activity points at your own risk configuration.
Treat "do not honour" with suspicion. It is a catch-all with no diagnostic content, and a rise in it is a signal to look at what changed around those attempts, not an explanation in itself.
Statistics locate the segment. Orders explain it. Order search accepts session status, merchant, processor, gate, endpoint, project, card type and currency, as well as direct identifiers — order id, merchant invoice number, customer id, amount — and returns summaries with no cardholder PII. Order detail adds status, decline reason, routing path, the step-by-step processing trail, card metadata and a masked contact.
Find declined orders for <merchant> on <processor> between <dates>, <issuer country>, <card type>. Take ten of them and, for each, show status, decline reason, the route taken and the processing trail. Tell me what the ten have in common and where they differ.
The trail is the part worth reading closely. It shows at which stage the attempt stopped, whether the payment cascaded to a second provider or was tried once, and whether the same order failed repeatedly. A recurring trait across several orders is worth investigating, but the sample must be large and representative enough for the affected traffic segment before it is reported as a pattern.
Payment behaviour follows configuration, so the route carrying the declines deserves a direct read. The pattern is search, resolve, read: find the record, resolve the name to an id, read the fields.
Show the endpoint, gate and processor records used by these orders. For the endpoint: payment strategy, capture and return timeouts, form templates and flags. For the gate: processor, currency, rate plan. For the processor: type, code, status. Compare them with the equivalent records on the route where the approval rate did not move.
The honest limit sits here, and it is worth stating plainly. MCP returns the configuration as it stands now — not as it stood last Tuesday. It cannot tell you what changed or who changed it. Setting today's live values against your own change record is a human step, and in a decline investigation it is usually the step that closes the case.
The characteristic failure of an AI-assisted investigation is a confident story assembled from one query. The correction is structural: make the model file its output in three bins and forbid it from smoothing over the third.
Summarise this investigation in three sections. FACTS - only statements you can support with a specific MCP result, naming the tool and the filters used. HYPOTHESES - plausible explanations not supported by data you can read. For each, state the check that would confirm or reject it. NOT AVAILABLE - questions I asked that this MCP access cannot answer. Do not merge the sections. Do not fill gaps with assumptions.
A short NOT AVAILABLE section means the investigation was well scoped. A long one is not a failure either — it tells you exactly which system to open next and which colleague owns it.
A finding is ready to pass on when someone else can act without re-running the analysis:
Then route it. Issuer-side reason clusters go to the processor or acquirer contact. A rise in fraud-filtered volume goes to the risk team as a threshold question. Timeout and trail anomalies go to engineering. A sudden shift in card type or issuer-country mix may reflect changes in checkout, campaigns, customer acquisition, geography or other traffic sources and should normally be checked with the merchant team.
Claude can query permitted operational data, but it cannot act. There is no write, update or execute operation on the MCP server, so no retry, refund, routing change or setting change can originate from the conversation. Those actions require separate authorised workflows outside MCP, such as the Payneteasy backoffice or the applicable APIs.
Nor does an investigation based on permitted operational data prove causation. What it produces is a narrowed population and a testable hypothesis — a faster starting point than manually cross-referencing separate dashboards, but still a hypothesis that needs the human steps above (configuration history, the relevant colleague, the final call) to close.
Thank you for reaching us. Your request has been sent successfully. We will get back to you as soon as possible.
Message was not sent