For those who read my blog, I tend to write about the future; the interesting technological and business changes on the horizon and how to prepare for them as IT leaders. Yet the topic of today’s short, but not-so-sweet article is about as old as I am: it’s enterprise monitoring.
Monitoring Apathy Always Equals Pain
Monitoring information systems isn’t cool or interesting to most people. That is, until systems go down and executives want answers. By no means am I exaggerating when I say I’ve seen my fair share of heads roll over the years for poor monitoring practices. Whether the IT department knows it or not, monitoring is a core competency it absolutely must possess.
I’ve heard every excuse in the book for why a system wasn’t monitored:
- It was a “shadow IT system” we didn’t know about!
- We thought (insert_other_department_here) was monitoring that system!
- We don’t have the budget for monitoring tools!
The list of excuses goes on. Yet the outcome is always the same: corporate rage. If you’re not monitoring critical systems, and not stepping up to manage outages, you’re going to feel the wrath of angry executives.
Your Job Security Hinges on Monitoring
In essence, monitoring and observability are not just technical necessities; they are strategic imperatives for modern enterprises. They enable organizations to navigate the complexities of the digital landscape, respond to challenges with agility, and maintain a resilient and high-performing IT infrastructure. As organizations continue to embrace digital transformation, the importance of comprehensive monitoring and observability will only grow, making them indispensable tools for IT leaders in today’s dynamic business environment.
A 2-Step Monitoring Get Well Plan
For those who’ve read this far, there’s good news: monitoring isn’t rocket science. It’s also a discipline that lends itself well to an iterative approach. That is, IT can start with a simple approach that creates value through visibility, then iterate and improve monitoring capabilities once the company sees the value of IT’s responsibility in cross-functional observability.
Step 1: Questions IT Leaders Must Ask Today
If you’re a CIO or infrastructure leader, there are three questions you need to ask of your senior infrastructure managers today:
- Have we cataloged critical assets across the organization?
- Are we monitoring them?
- Do we have procedures if these systems go down?
When I say ask these questions today, I mean so in the literal sense. This is information that you as an IT leader must know. If you don’t ask these questions yourself, a leader above you will do this job for you. I assure you that won’t be pretty.
Step 2: Create Your Systems Lifecycle 1.0 Capability
Assuming you have monitoring gaps in your organization, create a 60-day improvement plan. This involves five basic tasks:
- Create a simple definition of what critical means. Write a single sentence, such as: “A system that is used by more than 20% of the company, a system that generates revenue, or a system that could impede core business operations if it goes down for more than one hour.”
- Catalog known systems. Using the definition from above, task 2-5 people with cataloging systems that fit this criteria. Blend a bottom-up and top-down approach. For the top-down approach: have a project manager quickly survey business unit leaders. For the bottom-up approach: have IT veterans jot down the systems they know of. This isn’t an exact science; the intent is to iterate and move with speed and urgency.
- Assign owners to systems. Don’t waste time obsessing over business owners versus system owners. Focus on who to notify if/when a system goes down.
- Start monitoring. Create basic up/down monitoring for key systems. If you have advanced application performance monitoring (APM) or other sophisticated observability tooling, great. Otherwise, perform the minimum viable monitoring to yield a logical representation of system availability.
- Operationalize incident management. When critical systems go down, initiate a basic incident management process. At minimum, this entails notifying the system owner, creating communications messaging during the incident, working with the vendor or in-house engineers to resolve the issue, and finally creating a root cause analysis (RCA) after the incident has been resolved.
A Closing Note on Real World Complexities
In real-world companies, there will be plenty of systems IT doesn’t fully manage. There will be business unit leaders who don’t like IT inserting itself into a position of responsibility. To that, I simply say: who cares!
As long as monitoring doesn’t knock over systems, it can be done with zero system impact and very few resources. Include business leaders in the monitoring process, but don’t ask for permission. Simply inform these leaders that monitoring will be done, and they will see value in the process.