NEWTickerr MCP is live →
tickerr
Swarmsourced · Tickerr
World's first swarmsourced LLM status platform
Major OutageResolvedOfficial report

Gemini Major Outage - July 25, 2026

Provider incident title: “RESOLVED: Google Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes, and Bare Metal Solutions (BMS) services are experiencing a service outage in europe-west4-a due to a cooling failure.

Share on X

Started

Saturday, July 25, 2026

01:16 PM UTC

Duration

< 1 minute

total

Resolved

01:16 PM UTC

Jul 25, 2026

Severity: Major Outage|Affected: Gemini|Duration: < 1 minute

Incident Timeline

01:16 PM UTCResolved

&lt;p&gt; Incident began at &lt;strong&gt;2026-07-15 16:57&lt;/strong&gt; and ended at &lt;strong&gt;2026-07-16 05:25&lt;/strong&gt; &lt;span&gt;(all times are &lt;strong&gt;US/Pacific&lt;/strong&gt;).&lt;/span&gt;&lt;/p&gt;&lt;div class="cBIRi14aVDP__status-update-text"&gt;&lt;h2&gt;# Incident Report&lt;/h2&gt; &lt;h2&gt;## Summary&lt;/h2&gt; &lt;p&gt;On Wednesday, 15 July 2026, Google Cloud VMware Engine (GCVE), Bare Metal Solution (BMS), and Google Cloud NetApp Volumes (GCNV) experienced service interruptions for a total duration of 14 hours, 55 minutes.&lt;/p&gt; &lt;p&gt;The root cause for this outage is a 3ms voltage drop in the power feed from the utility provider and subsequent utility breaker protective action. This incident was mitigated by the data center provider rectifying the failed systems followed by the Google engineering team restoring the services.&lt;/p&gt; &lt;h2&gt;## Root Cause&lt;/h2&gt; &lt;p&gt;A regional data center hosting services in europe-west4-a experienced an upstream voltage transient affecting both utility power feeds A and B. During this event, the utility breakers on both A and B feeds tripped and initiated the transfer to the back-up power source DRUPS (Diesel Rotary Uninterruptible Power Supply). The transition of side B to DRUPS system was successful without any power interruption. The DRUPS back-up power system for side A failed to take over the facility load due to electrical component failures.&lt;/p&gt; &lt;p&gt;The DRUPS failure to take over load initiated automatic transfer of the affected 3 rows from feed A to redundant power feed B. Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and experienced complete power loss due to an overload protection breaker trip. The root cause of the failed transfer on the affected row was identified as load deployment discrepancy which is under further investigation by Google engineering team. This resulted in loss of both redundant power feeds to the single row.&lt;/p&gt; &lt;p&gt;During the voltage transient event, the server data hall experienced an increased temperature due to a cooling system failure. The chiller controller dropped offline during the voltage transient event, failing to signal the chilled water distribution pumps to restart and ultimately causing the chiller system A to shut down. The redundant source was not available due to known ongoing construction work at the facility. This resulted in the data hall temperatures reaching 44°C and subsequent shutdown of the affected data hall machines. Google engineering teams initiated machine shutdown procedures for the remainder of reachable devices as part of the cooling emergency shutdown process.&lt;/p&gt; &lt;p&gt;A timeline of events during the incident is provided below.&lt;/p&gt; &lt;h3&gt;DataCenter Events&lt;/h3&gt; &lt;table&gt; &lt;thead&gt; &lt;tr&gt; &lt;th align="left"&gt;Timestamp(PST)&lt;/th&gt; &lt;th align="left"&gt;Event Description&lt;/th&gt; &lt;/tr&gt; &lt;/thead&gt; &lt;tbody&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:24&lt;/td&gt; &lt;td align="left"&gt;An electrical fault occurred on the utility grid upstream of the data center, disrupting the electrical distribution.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:24&lt;/td&gt; &lt;td align="left"&gt;DRUPS back-up power system for side A failed to take over the data hall load.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:24&lt;/td&gt; &lt;td align="left"&gt;The chiller controller dropped offline during the voltage transient event causing distribution pumps to stop and unable to restart.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:37&lt;/td&gt; &lt;td align="left"&gt;Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and lost power.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 18:29&lt;/td&gt; &lt;td align="left"&gt;Data hall temperatures reached 44C and crossed the safe machine operating threshold.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 19:55&lt;/td&gt; &lt;td align="left"&gt;Machine shutdown procedures implemented.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:05&lt;/td&gt; &lt;td align="left"&gt;Cooling system fully recovered and temperatures in the data hall returned to normal operating range.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:24&lt;/td&gt; &lt;td align="left"&gt;Notification about cooling recovery sent by provider&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:46&lt;/td&gt; &lt;td align="left"&gt;Notification about power recovery sent by provider&lt;/td&gt; &lt;/tr&gt; &lt;/tbody&gt; &lt;/table&gt; &lt;h3&gt;GCVE Timelines&lt;/h3&gt; &lt;table&gt; &lt;thead&gt; &lt;tr&gt; &lt;th align="left"&gt;Timestamp(PST)&lt;/th&gt; &lt;th align="left"&gt;Event Description&lt;/th&gt; &lt;/tr&gt; &lt;/thead&gt; &lt;tbody&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:30&lt;/td&gt; &lt;td align="left"&gt;GCVE service power monitoring alert received&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:41&lt;/td&gt; &lt;td align="left"&gt;First symptom detected, servers reporting redundancy power feed failure.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 17:05&lt;/td&gt; &lt;td align="left"&gt;Switch temperature &amp;gt;60°C alerts triggered.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 17:09&lt;/td&gt; &lt;td align="left"&gt;Start of user impact, first prober alert failure for north south traffic for customers.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 19:00&lt;/td&gt; &lt;td align="left"&gt;Server / Network devices shutdown initiated.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 19:55&lt;/td&gt; &lt;td align="left"&gt;All reachable server/network devices were shut down.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:24&lt;/td&gt; &lt;td align="left"&gt;Cooling system fully recovered and temperatures in the datahall were returning to normal operating range.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:56&lt;/td&gt; &lt;td align="left"&gt;Network restoration started with reachable hydra devices.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 22:40&lt;/td&gt; &lt;td align="left"&gt;Onsite technician arrived, to recover the console servers.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 23:10&lt;/td&gt; &lt;td align="left"&gt;Console server reboot/recovery complete.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-16-2026 00:33&lt;/td&gt; &lt;td align="left"&gt;Network recovery for all Placement Groups complete.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-16-2026 02:36&lt;/td&gt; &lt;td align="left"&gt;Incident mitigated, after recovering all the customer PCs are fully healthy.&lt;/td&gt; &lt;/tr&gt; &lt;/tbody&gt; &lt;/table&gt; &lt;p&gt;This outage impacted 24 private clouds of 20 distinct customers in europe-west4-a region.&lt;/p&gt; &lt;h3&gt;Bare Metal Solutions (BMS) Timelines&lt;/h3&gt; &lt;table&gt; &lt;thead&gt; &lt;tr&gt; &lt;th align="left"&gt;Timestamp(PST)&lt;/th&gt; &lt;th align="left"&gt;Event Description&lt;/th&gt; &lt;/tr&gt; &lt;/thead&gt; &lt;tbody&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:41&lt;/td&gt; &lt;td align="left"&gt;First symptom detected, servers and storage reporting redundancy power feed failure&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 17:52&lt;/td&gt; &lt;td align="left"&gt;Automated monitoring detected rising temperatures&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 18:37&lt;/td&gt; &lt;td align="left"&gt;Four Netapp SAN storage nodes failed and were shut down due to overheating, resulting in storage availability issue&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 18:37&lt;/td&gt; &lt;td align="left"&gt;Customer impact started, first server shutdown detected&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 18:59&lt;/td&gt; &lt;td align="left"&gt;Reserve and buffer servers started to get powered off&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 20:43&lt;/td&gt; &lt;td align="left"&gt;Drop in the temperature has been detected&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 20:52&lt;/td&gt; &lt;td align="left"&gt;Four Netapp SAN storage nodes recovered, mitigating storage availability issue&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:05&lt;/td&gt; &lt;td align="left"&gt;The cooling system fully recovered and temperatures in the datahall were returning to normal operating range.&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 23:23&lt;/td&gt; &lt;td align="left"&gt;First customer server reboot started&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-16-2026 06:32&lt;/td&gt; &lt;td align="left"&gt;Last customer server reboot started&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-16-2026 08:42&lt;/td&gt; &lt;td align="left"&gt;Incident mitigated. All impacted customer servers were rebooted and confirmed in healthy state.&lt;/td&gt; &lt;/tr&gt; &lt;/tbody&gt; &lt;/table&gt; &lt;p&gt;This outage impacted 9 distinct BMS customers in europe-west4&lt;/p&gt; &lt;h3&gt;NetApp Timelines&lt;/h3&gt; &lt;table&gt; &lt;thead&gt; &lt;tr&gt; &lt;th align="left"&gt;Timestamp(PST)&lt;/th&gt; &lt;th align="left"&gt;Event Description&lt;/th&gt; &lt;/tr&gt; &lt;/thead&gt; &lt;tbody&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 16:32&lt;/td&gt; &lt;td align="left"&gt;Netapp facility remote monitoring alert on high temp and network switches switches failure received&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 17:32&lt;/td&gt; &lt;td align="left"&gt;All 6 clusters failed and were shut down automatically due to high temperature&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 21:05&lt;/td&gt; &lt;td align="left"&gt;Cooling capacity restored&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 22:08&lt;/td&gt; &lt;td align="left"&gt;Cluster recovery process initiated for all 6 clusters. 1st cluster recovered&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td align="left"&gt;07-15-2026 23:55&lt;/td&gt; &lt;td align="left"&gt;All 6 Netapp Clusters and all customers recovered and validated as healthy &amp;amp; operational&lt;/td&gt; &lt;/tr&gt; &lt;/tbody&gt; &lt;/table&gt; &lt;h2&gt;## Remediation and Prevention&lt;/h2&gt; &lt;p&gt;Google engineers were alerted to the infrastructure power loss on Wednesday, 15 July 2026 at 16:30 US/Pacific and subsequent increase in data hall temperatures at 17:44 US/Pacific via our automated monitoring and telemetry, and immediately began investigating.&lt;/p&gt; &lt;p&gt;During the voltage transient event, the chillers maintained power but shut down because the distribution pumps failed to restart, resulting in no water flow. The pumps were manually switched from auto to hand mode to restore circulation and cooling. The facility team deployed an interim portable UPS to support the chiller controllers and prevent localized controller power loss.&lt;/p&gt; &lt;p&gt;Upon confirmation of utility power restoration, the facility returned to utility power via automatic switching with the exception of two Remote Power Panels (RPPs), which required manual verification before restoration of power to affected server row. To provide full resiliency the facility team aligned a swing DRUPS unit to restore the redundant configuration on the A feed and is working on repairing the faulty DRUPS.&lt;/p&gt; &lt;p&gt;To remediate the power loss at Row 3, once the electrical fault was confirmed to be fully isolated, teams reset the tripped breakers, successfully restoring both redundant power to the impacted server racks.&lt;/p&gt; &lt;p&gt;Google is committed to preventing a repeat of this issue in the future and is completing the following actions:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;Detailed investigation by the engineering team of the sequence of events that prevented transfer of load to DRUPS system and identified system improvements &lt;strong&gt;[ETA Aug 2026].&lt;/strong&gt;&lt;/li&gt; &lt;li&gt;Determine and resolve the cause of overload conditions that caused power loss to Row 3 &lt;strong&gt;[ETA Aug 2026].&lt;/strong&gt;&lt;/li&gt; &lt;li&gt;Perform investigation of chiller pump control system redundancy setup to ensure cooling system resiliency &lt;strong&gt;[ETA Sep 2026].&lt;/strong&gt;&lt;/li&gt; &lt;li&gt;Implement improvements to facility system monitoring alerts for power events and configure early alert thresholds for cooling excursion detection. &lt;strong&gt;[ETA Aug 2026].&lt;/strong&gt;&lt;/li&gt; &lt;li&gt;The GCVE engineering team is working to improve the automated triggering of shutdown during quickly evolving thermal runoff situations &lt;strong&gt;[ETA Oct 2026]&lt;/strong&gt;.&lt;/li&gt; &lt;li&gt;The BMS engineering team is working with partner teams to update the incident classification and SLO definitions to improve personnel availability during major datacenter events &lt;strong&gt;[ETA Sep 2026]&lt;/strong&gt;.&lt;/li&gt; &lt;li&gt;The GCNV engineering team is collaborate with the data center team to review incident handling SLA and runbooks, identify potential future occurrences and create playbooks to address/mitigate them, establish monthly joint emergency drill, and review if the Netapp cluster recovery process can be expedited &lt;strong&gt;[ETA Aug 2026].&lt;/strong&gt;&lt;/li&gt; &lt;/ul&gt; &lt;h2&gt;## Detailed Description of Impact&lt;/h2&gt; &lt;p&gt;On 15 July 2026 from 16:39 to 16 July 2026 07:34 US/Pacific, customers experienced service disruptions across several products in europe-west4-a zone:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;&lt;strong&gt;Google Cloud VMware Engine&lt;/strong&gt;: Customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection.&lt;/li&gt; &lt;li&gt;&lt;strong&gt;Bare Metal Solution&lt;/strong&gt;: Customers experienced connectivity loss to their database servers and storage appliances as the underlying hardware was powered down.&lt;/li&gt; &lt;li&gt;&lt;strong&gt;Google Cloud NetApp Volumes&lt;/strong&gt;: Customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, experienced failures in the affected region.&lt;/li&gt; &lt;/ul&gt; &lt;/div&gt;&lt;hr&gt;&lt;p&gt;Affected products: Bare Metal Solution, Google Cloud NetApp Volumes, VMWare engine&lt;/p&gt;&lt;p&gt;Affected locations: Netherlands (europe-west4)&lt;/p&gt;

Get alerted next time

Install Tickerr MCP — your agent auto-reports & routes around outages

When Gemini goes down again, your agent reports anonymously and instantly gets a fallback recommendation from the swarm.

$claude mcp add tickerr --transport http --url https://tickerr.ai/mcp
Install free →

Gemini is back online

View current status and uptime history

Gemini status →

About This Report

Gemini 30-day uptime: 0% based on Tickerr's independent monitoring checks.

This incident was sourced from Gemini's official status RSS feed. Tickerr polls RSS feeds every 10 minutes and merges updates from the same outage into a single incident timeline.

This incident lasted < 1 minute.

Tickerr monitors 90+ AI tools independently. View live Gemini status or all AI tool status.

Other Gemini incidents

Jul 24, 2026
Jul 20, 2026
Jul 18, 2026
Jul 16, 2026
Degraded Performance

gemini-2.5-flash-lite API Latency Degraded

Jul 15, 2026

View all Gemini incident history →

Get alerted next time Gemini goes down

Tickerr monitors 90+ AI tools. We'll email you when an incident starts or resolves.

Weekly AI pricing & uptime digest

Price drops, new model releases, and incident summaries - every Monday. Free.

Also on Tickerr

Gemini live statusGemini pricingAll AI tool statusAI model pricingCompare AI toolsAI updates