Intended Audience:
- Drupal developers, administrators and consultants with a general knowledge of Drupal.
- Anyone interested in Drupal's data model.
- Anyone interested in the landscape of available Drupal data cleanup tools.
Topics covered:
- Drupal's handling of node revisions.
- Pitfalls of being too clever when implementing custom code in Drupal.
- Available data cleanup solutions and what was missing from each one.
- How we achieved all 3 of our main goals: performance, accuracy and resilience.
Summary
Our story begins innocently enough, when a colleague mentions that they're having trouble creating a local database backup. What ensues is a weeks-long hunt for answers across space, time, and Drupal.org.
Every time we think we've figured it out, The Blob just keeps growing! Follow our intrepid team of Drupalistas as we explore new worlds of previous consultants' code and "loosely documented" modules on Drupal.org to reveal that the function calls were coming from inside the house all along (dramatic space opera music).
You'll laugh, you'll cry, you'll be on the edge of your seat, or at least, you'll learn a thing or two about handling millions of redundant node revisions.
Tune in to find the answers to these questions and more:
- Will our crack team of heroes succeed in stopping The Blob from growing before it exceeds the size of the known universe?
How we figured out the source of the problem and stopped it. - Even if they do, how will they clean up the galactic mess it left in its wake?
Deleting millions of rows of data turns out to be a technically complex and demanding task! - Will they manage to delete only the bogus data, while keeping the precious valid data safe?
How could we tell that we were deleting only the bad data? - Will they make it home in time for dinner?
Since it turned out to take a very long time, how could the cleanup process be safely interrupted and resumed the next day?