TARGET Services Outage Case Study

TARGET Services Outage Case Study

On 27 February 2025, TARGET Services suffered a major operational incident affecting the infrastructure through which central-bank money, securities and collateral move across Europe.

TARGET Services comprise:

  • T2, which processes high-value euro payments;
  • TARGET2-Securities, which provides securities settlement;
  • TARGET Instant Payment Settlement, which supports instant payments; and
  • The Eurosystem Collateral Management System, introduced later in 2025.

Shortly after 08:00, communication problems began affecting TARGET2-Securities. Support teams discovered a memory shortage in a database buffer pool used to process messages and transactions. The settlement engine continued processing instructions already received, but new messages could not flow normally.

The problem later affected T2. Initial evidence suggested a database or software malfunction. Because data and software conditions could be replicated to the secondary site, teams initially believed that an immediate failover might reproduce the problem instead of resolving it.

The investigation eventually identified a hardware failure involving a core component of the storage system. According to the European Central Bank’s post-mortem, two redundant elements within that component failed simultaneously. The equipment vendor described the behaviour as something it had not previously encountered anywhere in the world.

Once the hardware cause was established, the Eurosystem initiated recovery to a secondary site. Additional data-consistency and integrity checks were required before processing could safely resume. Consequently, the contractually agreed technical recovery-time objective of 60 minutes was not achieved.

T2 remained unavailable for approximately 10 hours, while TARGET2-Securities was unavailable for around eight hours. TARGET Instant Payment Settlement experienced partial disruption for approximately one hour. Nevertheless, extended operating hours allowed most payment activity to be completed by the end of the day. ECB TARGET Services post-mortem.

The incident demonstrates that operational resilience is determined by more than duplicate hardware and a secondary site. Effective resilience also requires rapid fault diagnosis, monitoring capable of detecting rare component behaviour, clear failover authority, coordinated communication and contingency arrangements covering every connected market infrastructure.

Why TARGET Services are systemically important

TARGET Services are operated by the Eurosystem to enable the movement of cash, securities and collateral across Europe. Transactions settle in central-bank money, limiting settlement exposure between participating institutions. European Central Bank overview of TARGET Services.

T2 supports:

  • Monetary-policy operations;
  • Bank-to-bank payments;
  • Commercial payments;
  • Settlement by ancillary payment systems;
  • Liquidity transfers; and
  • Cash settlement for other financial-market infrastructures.

TARGET2-Securities provides a common platform for securities settlement. Its availability affects central securities depositories, banks, investment firms and other market participants.

These services are interconnected. A problem affecting a shared infrastructure component can therefore spread beyond its apparent technical origin.

The scale is significant. During 2025, T2 settled euro payments valued at approximately €492.9 trillion. Urgent payments represented more than one-fifth of total turnover by value. TARGET Services Annual Report 2025.

An interruption can consequently affect liquidity, securities delivery, retail-payment-system settlement and foreign-exchange arrangements, even when individual commercial banks remain technically operational.

What happened?

Initial disruption

Shortly after 08:00 on 27 February 2025, a disruption was detected in application-to-application communication between TARGET2-Securities and its network service providers.

A memory shortage affected a database group buffer pool used to cache information required for message and transaction processing. Although instructions already present in the settlement engine continued to be processed, incoming communications could not be handled normally.

Attempts to resolve the problem by allocating additional space were unsuccessful. The service desk closed affected communication channels to reduce the number of undeliverable messages.

Spread to T2

Shortly after 10:00, similar problems began affecting both inbound and outbound T2 messages.

The two services were affected at different times because their database buffer pools were segregated. However, the underlying hardware problem ultimately caused both pools to become saturated.

This sequence illustrates an important concentration issue: logical separation between applications does not ensure resilience when they depend on a common physical infrastructure layer.

Initial diagnosis

Support teams initially suspected a database or software problem. That diagnosis affected the recovery decision.

If the problem had existed within synchronised software or data, transferring operations to the secondary site might have replicated the failure. Teams therefore first attempted shutdowns, selective restarts and additional investigation.

The measures increased available space temporarily but did not remove the underlying problem.

Hardware failure identified

Technical specialists and the responsible external vendor subsequently identified faulty behaviour in a critical component of the storage system.

The ECB reported that the underlying event involved the unexpected simultaneous failure of two elements intended to provide redundancy. The vendor stated that similar behaviour had not previously been detected in comparable components worldwide.

The component required physical replacement. However, immediate service recovery was pursued through a transfer to the secondary production site rather than waiting for physical repair.

Site recovery

Site-recovery activities began at 15:35 after the relevant crisis managers approved the decision.

The standard recovery procedure envisaged a technical recovery time of one hour. In this case, the nature of the disk-storage failure required additional checks to establish:

  • Data consistency;
  • Data integrity;
  • Proper functioning of storage at the secondary site; and
  • The ability to restart the business day safely.

Operations resumed shortly after 18:00. The faulty physical component was replaced on 28 February.

Verified timeline

Time/date Development
Shortly after 08:00, 27 February Communication disruption detected in TARGET2-Securities
Morning Database buffer-pool memory shortage identified
Morning Communication channels closed to limit undeliverable messages
Shortly after 10:00 Similar processing problems began affecting T2
Early afternoon Shutdown and selective-restart attempts did not resolve the issue
Later afternoon Hardware storage-system failure identified with vendor assistance
15:35 Site recovery initiated after crisis-management approval
Shortly after 18:00 Operations resumed from the secondary site
By midnight Services completed an orderly close of the operating day
28 February Faulty component replaced
March 2025 European Securities and Markets Authority addressed settlement penalties connected with the outage
2025 Twenty follow-up measures progressed through TARGET governance
November 2025 Detailed public post-mortem issued

Business and financial-market impact

Payments

When the incident began, T2 had processed approximately:

  • 55% of the day’s payment volume; and
  • 35% of its expected payment value.

Operating hours were extended after recovery. As a result, the number and total value of payments settled that day were broadly comparable with previous days.

However, some national central banks reported possible delays in salary and pension payments. Although the relevant retail systems completed interbank settlement later that afternoon, certain institutions did not have sufficient time to credit beneficiaries before the end of their internal processing day. Those transfers were completed on the following calendar day.

Connected market infrastructures

The outage produced spillover effects:
  • EURO1 and STEP2 settlement was postponed;
  • CLS payouts in several currencies were suspended for hours;
  • Liquidity transfers between TARGET services were interrupted;
  • Some ancillary systems could not settle at their normal times; and
  • Central counterparties used contingency arrangements for certain margin payments.

EURO1 and STEP2 eventually settled successfully without rejected payment files.

Securities settlement

TARGET2-Securities processed a volume of transactions broadly comparable with normal days, but efficiency declined.

Settlement efficiency by volume fell to 90.1%, compared with a 2024 average of 94.4%. By value, efficiency fell to 83.5%, compared with a 2024 average of 97.7%.

Penalties associated with late matching and failed settlement increased substantially. The European Securities and Markets Authority subsequently stated that cash penalties with detection dates of 27 and 28 February 2025 should not apply.

Wider financial stability

Despite the severity of the interruption:

  • Short-term euro interest rates were not materially affected;
  • The euro short-term rate remained close to previous levels;
  • Most payments were completed after restoration;
  • No disorderly end-of-day failure was reported; and
  • Market confidence was not materially destabilised.

This outcome should not obscure the risk. It demonstrates that extended operating hours, participant cooperation and contingency arrangements successfully prevented a technical failure from developing into a wider liquidity or financial-stability incident.

Risk and control-failure analysis

1. Common-component concentration risk

T2 and TARGET2-Securities were logically separated but depended on a common type of physical storage infrastructure.

The case demonstrates three different forms of concentration:

  • Physical concentration: Multiple services use common hardware layers.
  • Provider concentration: Diagnosis depended partly on one specialised external vendor.
  • Process concentration: Payments, securities and liquidity transfers relied on closely connected TARGET components.

Risk maps should therefore extend below business applications to storage, databases, networks, power, identity systems and specialised vendor support.

2. Redundancy-design risk

The storage component contained redundant elements, but two of them failed simultaneously.

Redundancy is effective only when:

  • Failure modes are sufficiently independent;
  • Monitoring detects deterioration before total failure;
  • Duplicate components do not share an unidentified weakness;
  • Failover does not depend on the same control plane; and
  • Recovery arrangements are periodically tested against common-cause scenarios.

Boards should not accept “the system is redundant” as sufficient assurance. They need evidence about the failure modes against which the redundancy has been designed.

3. Diagnostic and observability risk

Several hours were required to determine whether the problem originated in the database, software or hardware.

Initial symptoms appeared at the database and messaging levels even though the root cause was a storage component. This delayed the failover decision because recovery actions appropriate for one cause could have been ineffective or dangerous for another.

Monitoring should correlate:

  • Hardware telemetry;
  • Storage performance;
  • Database-buffer behaviour;
  • Message queues;
  • Application latency;
  • Capacity utilisation; and
  • Error patterns across interconnected services.

The vendor introduced additional monitoring and a dedicated process for handling recurrence of the rare hardware behaviour.

4. Recovery-time risk

The contractual technical recovery-time objective was 60 minutes, but the actual disruption lasted much longer.

The one-hour measure applied to the failover procedure itself, not the entire period from service disruption to recovery. Time was also consumed by diagnosis, approvals and integrity checks.

Organisations should therefore distinguish between:

  • Time to detect;
  • Time to diagnose;
  • Time to decide;
  • Time to initiate recovery;
  • Technical failover time; and
  • Time to restore the complete business service.

A recovery-time objective that begins only after a formal failover decision can materially understate customer and market impact.

5. Data-integrity risk

Rapid restoration cannot come at the expense of settlement finality or data integrity.

Because the incident involved storage infrastructure, teams needed assurance that payment and securities records were consistent before restarting. An incorrect restart could have created duplicate instructions, missing transactions or reconciliation disputes.

For core financial infrastructure, the recovery decision must balance speed against certainty that books and settlement positions remain accurate.

6. Crisis-governance risk

Separate T2 and TARGET2-Securities crisis-management calls delayed some joint decisions. The post-mortem concluded that combined calls should be considered when the same event affects both services.

Crisis governance should define:

  • Who can order a failover;
  • What evidence is necessary;
  • When joint governance is activated;
  • How disagreements are escalated;
  • Which integrity checks are mandatory; and
  • Who can extend settlement cut-off times.
7. Communication risk

The ECB website provided roughly hourly updates, and market participants generally recognised the increased transparency.

However, weaknesses remained:

  • Some securities participants misunderstood communications about postponed penalty calculations;
  • T2 participants wanted earlier instructions to stop sending messages;
  • The T2 crisis-communication group was activated, but the corresponding TARGET2-Securities group was not;
  • Participants wanted advance notice of crisis calls; and
  • Communication asymmetry emerged between affected services.

The case shows that communication must be operationally actionable, not merely frequent.

8. Contingency-liquidity risk

When normal T2 processing is unavailable, contingency arrangements must prioritise critical payments. Participants also need alternative sources of liquidity when securities normally used for collateral cannot move through TARGET2-Securities.

Recovery testing should therefore include a combined scenario in which:

  • The payment system is unavailable;
  • The securities platform is unavailable;
  • Collateral cannot be transferred normally; and
  • Numerous participants need contingency access simultaneously.
Remediation and follow-up

TARGET governance identified 20 follow-up measures. Most had been implemented by the fourth quarter of 2025. They included:

  • Replacement of the faulty physical component;
  • Enhanced hardware monitoring;
  • Review of business-continuity and information-technology service-continuity arrangements;
  • Improved site-recovery procedures;
  • Greater capacity for contingency user sessions;
  • Testing scenarios in which TARGET2-Securities is unavailable;
  • Review of alternative liquidity sources;
  • Clearer definitions of critical transactions;
  • Better guidance on extending cut-off times;
  • Joint crisis-call arrangements;
  • Improved participant communication; and
  • Work on standard treatment of penalties after major incidents.

The ECB’s 2025 Annual Report confirmed that the incident and business-continuity processes were reviewed and that the identified measures were substantially progressed. ECB Annual Report 2025.

Current regulatory relevance

The Principles for Financial Market Infrastructures require systemically important infrastructures to identify operational risks, remove single points of failure and maintain robust business-continuity arrangements. CPMI-IOSCO Principles for Financial Market Infrastructures.

The case is also relevant to Europe’s Digital Operational Resilience Act, which has applied since 17 January 2025. The regulation requires covered financial entities to manage information-and-communications-technology risk, report significant incidents, conduct resilience testing and oversee third-party technology dependencies. Central banks themselves are exempt from DORA, but banks and other financial entities using critical infrastructures remain responsible for their own continuity arrangements. EUR-Lex overview of DORA.

For Indian banks and payment-system operators, the incident reinforces the importance of:

  • Mapping shared technology dependencies;
  • Testing primary and secondary sites;
  • Maintaining alternative payment arrangements;
  • Monitoring infrastructure capacity and component health;
  • Escalating major incidents promptly; and
  • Measuring resilience at the level of the complete payment service.

Lessons for boards and risk professionals

  1. Map infrastructure below the application level. Storage, networks, database components and vendor control systems can connect services that appear operationally separate.
  2. Challenge claims of redundancy. Boards should ask which common-cause failures could disable multiple redundant components.
  3. Measure total recovery time. Detection, diagnosis and decision delays must be included—not only the technical failover period.
  4. Test uncertain-cause scenarios. Recovery teams should practise situations where they cannot immediately determine whether a failure is caused by hardware, software, data or connectivity.
  5. Protect data integrity during recovery. Fast restart is unacceptable if payments or securities records cannot be reconciled confidently.
  6. Coordinate interconnected services. Joint incidents require joint crisis governance and communication.
  7. Prepare participants, not only operators. Banks and other users need their own liquidity, reconciliation and customer-communication procedures.
  8. Treat communication as a control. Updates must tell participants what to stop, continue or prioritise.
  9. Test market-wide capacity. Contingency facilities that work for one institution may fail when many participants need them simultaneously.
  10. Learn from near-systemic events. The absence of financial instability does not mean the original controls were adequate.

Practical financial-infrastructure resilience checklist

  • Are all critical payment and settlement services formally identified?
  • Are shared physical and logical dependencies mapped?
  • Can two redundant components fail through the same mechanism?
  • Does monitoring cover hardware, storage, database and message-processing layers?
  • Can teams distinguish hardware failure from replicated software or data corruption?
  • Is total service-restoration time measured from the first disruption?
  • Can primary and secondary sites be operated independently?
  • Are data-integrity checks predefined and automated where possible?
  • Can failover authority be exercised without sequential governance delays?
  • Are joint crisis calls available for interconnected services?
  • Can participants obtain liquidity if the normal collateral platform is unavailable?
  • Are contingency systems sized for simultaneous market-wide demand?
  • Can cut-off times be extended through a predefined decision process?
  • Do communications contain specific operational instructions?
  • Are penalty and compensation arrangements defined before an incident?
  • Are external vendors required to provide immediate specialist support?
  • Are corrective measures independently validated before closure?

Key takeaways

The TARGET Services outage was not a cyberattack, global banking collapse or defective software update. It was a rare hardware failure within systemically important European financial infrastructure.

Its significance came from the simultaneous interruption of payments, securities settlement and liquidity transfers. Logical separation and redundant hardware did not prevent disruption because an unforeseen common component failure affected multiple services.

Most payment activity was ultimately completed, and wider financial stability was preserved. That successful outcome reflected extended operating hours, contingency measures and coordinated recovery—not an absence of serious risk.

The principal lesson is that operational resilience must be assessed end to end. Organisations must understand shared infrastructure, challenge assumptions about redundancy, diagnose failures rapidly, coordinate recovery decisions and prepare for circumstances that existing technical models classify as extremely unlikely.

author avatar
RMA INDIA

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.