Use cases
Software Products E-commerce MSPs Schools Development & Marketing DevOps Agencies Help Desk
Company
Internet Status Blog Pricing Log in Get started free

Outage in AgResearch eRI

Tamaki Data Centre is overheating, all nodes potentially being shutdown

Resolved Minor
July 28, 2026 - Started 17 days ago - Lasted 6 days
Official incident page

Incident Report

We are currently investigating this issue.

Trusted by 1,000+ teams

The Status Page Aggregator with Early Outage Detection

Stop finding out about outages from your users. Monitor 6,320+ cloud services and get alerted the second something breaks.

IsDown status aggregator dashboard
Latest Updates ( sorted recent to last )
RESOLVED 11 days ago - at 08/03/2026 08:39PM

The primary incident has been resolved and a separate status incident is being created to reflect the current outages.

MONITORING 12 days ago - at 08/03/2026 02:04AM

Further investigation has shown that OnDemand is severely degraded/non operational. This is being investigated and will be resolved as soon as possible.

Access to and use of Globus also remains intermittent.

MONITORING 12 days ago - at 08/02/2026 10:03PM

All services that were resumed last week seem to be working as expected with all resources now available. All remaining systems have now been restarted.

Please note that REANNZ are unable to test systems where we don’t have access. Please don’t hesitate to contact support with any issue discovered.

MONITORING 15 days ago - at 07/31/2026 04:56AM

All but one hugemem nodes are up, and multiple users have had jobs start and finish as expected. Coldfront should now be operational. DMF is up and looks ok.

OnDemand is available but testing has shown degraded performance. We are continuing to work with our vendors to health check tape drives media.

Please note that REANNZ are unable to test systems where we don’t have access. Please don’t hesitate to contact support with any issue discovered.

We will resume providing updates on Monday.

IDENTIFIED 15 days ago - at 07/31/2026 03:10AM

We are looking into gpu and hugemem nodes, which continue to be unavailable at present.

Service restoration of Coldfront is ongoing but progressing well.

HPE engineers are currently helping investigating the health of DMF and tape media services.

We plan to provide our next update around 17:00.

IDENTIFIED 15 days ago - at 07/31/2026 12:57AM

We have brought OnDemand, login and all compute nodes back online and are currently monitoring their health status. Jobs can be submitted. GPFS/SMB access via S: and M: drives is available. All jobs that were running either failed or have been cancelled.

Service restoration efforts to Nix, GQuery and Coldfront are ongoing.

We expect to provide our next update around 15:00.

IDENTIFIED 15 days ago - at 07/30/2026 09:58PM

Access to eRI via ssh and OnDemand is still closed as we test the compute nodes. When the GPFS nodes are mounted, files will be available via Windows drive access.

We are bringing compute nodes back online but they will not be accessible until we have completed health checks on them. Nix, GQuery, Peaks and related VMs will all be brought online as part of that process.

Service restoration efforts are ongoing. We do not yet have enough information to provide an accurate recovery estimate, but we will continue to share regular updates. Unless the situation materially changes, we expect to provide our next update around 13:00.

IDENTIFIED 16 days ago - at 07/30/2026 05:06AM

Recovery of login services is progressing well, however this work is ongoing. The compute nodes are not online yet, batch jobs are not running.

We expect to provide an update at approximately 10:00am tomorrow.

IDENTIFIED 16 days ago - at 07/30/2026 02:59AM

HPC storage services appear to be in good health, although final checks are ongoing.

In parallel, we are working to get login services up and running however, at this stage we do not expect to have compute nodes online today.

Freezer and tape-based storage is expected to remain offline for longer, while we perform drive and media consistency checks. We will provide further information on this as we know more.

We plan to provide another update at approximately 17:00 today.

IDENTIFIED 16 days ago - at 07/30/2026 12:56AM

We have powered on essential hardware to begin bringing systems online. We are actively working on ramping up services starting with HPC storage and validating the status of storage. We are bringing up the minimum number of virtual instances to manage the rest of the fleet once storage is online.

We are keeping a very close eye on the environmental metrics at the data centre to prevent any additional fluctuations that might impact our tape media.

We plan to provide another update at approximately 15:00 today.

IDENTIFIED 16 days ago - at 07/29/2026 10:04PM

Tamaki Data Centre (TDC) has confirmed that the cooling system is back online and they are performing environmental checks etc.
We need to do a slow and controlled resumption of business as usual and perform health checks on the systems as we go.
We plan to provide another update about the progress and status by 13:00 today.

IDENTIFIED 16 days ago - at 07/29/2026 10:02PM

The current status is that we have the two working chillers at the data centre - TDC and a team is working on restoring the full chiller capacity.

IDENTIFIED 16 days ago - at 07/29/2026 05:22AM

Tamaki Data Centre has identified issues affecting two of the site's three chillers, resulting in reduced cooling capacity.
Teams are actively investigating the root cause and monitoring site temperatures and cooling performance. Our priority is the safe restoration of cooling services and maintaining the integrity of the platform.
Once temperatures have stabilised and cooling capacity has been validated, we will commence our recovery and startup plan. Further updates will be provided as the investigation progresses, this will likely be from tomorrow morning

INVESTIGATING 17 days ago - at 07/29/2026 05:02AM

We are continuing to investigate this issue.

INVESTIGATING 17 days ago - at 07/29/2026 04:24AM

Cooling at the Tamaki Data Centre has not yet been restored. All services and hardware remains down. We currently don't have any ETA.

INVESTIGATING 17 days ago - at 07/29/2026 03:21AM

Contractors are onsite at the TDC but have not yet confirmed the chiller issue

INVESTIGATING 17 days ago - at 07/29/2026 02:59AM

The data centre has chiller issues - Tamaki Data Centre (TDC) and all nodes are in danger of overheating. Consequently all nodes are being shutdown to protect the hardware

INVESTIGATING 17 days ago - at 07/29/2026 02:57AM

We are currently investigating this issue.

Latest AgResearch eRI outages

compute-3 "Not Responding - about 1 month ago
OnDemand is down - 3 months ago
Long waiter on compute-2 - 4 months ago

The Status Page Aggregator with Early Outage Detection

With IsDown, you can monitor all your critical services' official status pages from one centralized dashboard and receive instant alerts the moment an outage is detected. Say goodbye to constantly checking multiple sites for updates and stay ahead of outages with IsDown.

Start free trial

No credit card required · Cancel anytime · 6320 services available

Integrations with Slack Microsoft Teams Google Chat Datadog PagerDuty Zapier Discord Webhook