This is a article discussing survey results collected by NetForecast, appearing in Business Communications Review, May 2007 - found here. The basic theme is to survey how moving to APM Best Practices improved the overall IT management experience. I was a bit confused by the use of "benchmarking" with APM, which I would call the APM Pilot. But they do make a good definition of how process changes enhance the APM experience.
"The goal of APM best practices is to improve the performance outcome, and for the best outcome these best practices cannot stand alone. Each must be embedded into a continuous improvement process that ensures that application performance meets your business needs..."
They describe how "APM Benchmarking" will help an organization understand their capabilities and how they compare with other practitioners. This is what I call the 'Skills Assessment' (Chapter 3), and with other exercises for assessing enterpise visibility, lets you understand overall maturity of your APM practice. I created more of a idealized model, based on an amalgam of actual customer achievements. But I never had the opportunity to count how many of each customers achieved a particular level of APM maturity.
What is interesting that among those companies that participated, here is where they 'self-assessed':
The bulk of the population rated themselves on the low side and this is consistent with what I have been seeing over the last few years. Only 1 participant gave themselves a full rating, out of 329 total participants. Folks really need best practices to help them progress in using APM to address performance problems and show value for the new processes that will be established.
There is also a nice chart that summarizes the types of metrics that the participants are interested in.
Again, the chart title doesn't make sense but these are the metrics that the participants are interested in. The difference between what is important, and what is actually thresholded, confirms that visibility gaps exist and implementation maturity - which the authors define as "follow-through". I prefer to use an APM technology stack, which implies the same metrics, but separating network, server and transaction monitoring - which tool will get you those metrics.
Another interesting graph is the impediments to realizing an APM practice. THe authors call these "inhibitors" and it is a useful graph.
I explore these issues in Chapter 2 of my book s "entry points" and the situations which cause all of these varied ways to impede an APM initiative.
Monday, August 22, 2011
Monday, August 15, 2011
APM Best Practices Can Help Organizations Plan and Manage Cloud Migration
Here is the first installment of a short series of articles about using APM techniques to plan and realize Cloud Computing initiatives. Appearing in"Database Trends and Applications".
Sunday, July 3, 2011
Book tour begins
We decided to jam in a launch of the book tour {clients and workshops} before the full European holidays take effect. I am in Stockholm the last few days, which is nice... but my laptop crashed (fan failure) and I'm on a borrowed laptop, with a Swedish keyboard... and I am getting stressed. I have a deadline for a cloud article (mostly done) so I decided to goof off and update the blog. I will be in Denmark tomorrow Monday night and London on Wednesday night. Meetings so far are good but my 'spidy sense' is tingling. Might be jet lag... or an epiphany! I'll get some more data (visits) and digest.
An Interview on Application Performance Management with CA Technologies
Business Management Systems June 2011
A couple typos snuck through... no worries.
They do not track comments on the article... so feel free to opine here.
A couple typos snuck through... no worries.
They do not track comments on the article... so feel free to opine here.
APM best practices: A conversation with author Michael J. Sydor
Appearing in SearchSoftwareQuality.com 6 May 2011 Registration required {free}
Part 1 Introductory discussion
Part 2 Discussing staffing issues
Part 3 Discussing implementation
If you have not got the book yet... this gives you some flavor of the discussion...
They do not track comments on the article, so feel free to opine here!
Part 1 Introductory discussion
Part 2 Discussing staffing issues
Part 3 Discussing implementation
If you have not got the book yet... this gives you some flavor of the discussion...
They do not track comments on the article, so feel free to opine here!
Saturday, March 12, 2011
Artifacts for the Book
Seems I missed an upload during the final publishing push. While I sort that out with the publisher, here are the various docs that the book refers to, at Google Docs:
https://docs.google.com/leaf?id=0B9WYLZErvx39NzZlZjNjYTMtNjYyOS00NzcwLWIwMDQtZWYxNDE0ODYxNGE4&hl=en&authkey=CPjs-tEM
https://docs.google.com/leaf?id=0B9WYLZErvx39NzZlZjNjYTMtNjYyOS00NzcwLWIwMDQtZWYxNDE0ODYxNGE4&hl=en&authkey=CPjs-tEM
Friday, February 18, 2011
Incident Tracking and APM Maturity
I've being doing more Architecture Assessments (post-deployment) than Planning Assessments lately and have noticed something troubling. Folks in general have mature incident and trouble management practices for the applications that they operate/manage. Yet they do not do any tracking for any incidents related to the APM solution. What's going on?
I believe this is due to the overall misunderstanding about APM being 'just another monitoring tool'. Some folks think APM looks like many other tools - just with better capabilities. And they treat it just the same - a tool for operations to use when there is a problem; and back on the shelf when things are quiet. This ignores the 24x7 reality of the solution. We already know that this presumption leaves significant gaps in managing the capacity of the metrics storage component.
A small APM initiative can go for years before they run out of capacity and no one is managing this until the solution becomes unstable - and then they realize they have limited understanding. A growing APM initiative will run into this problem more quickly, depending on the pace of their successive deployments. But the pace of deployment is not as significant as much as the absence or presence of incident tracking for the APM solution.
In the absence of incident tracking, the client continues on blindly experiencing instability of the monitoring solution and then will escalate to vendor support, initialling labeling everything as a 'product defect'. The vendor support will then attempt to confirm the 'defect' but after finding nothing wrong (no known incidents, no prior history of instability), you end up with a bit of an impasse: no defect, yet no resolution because the problem is the configuration, not the monitoring software.
Why then is 'incident tracking' so magical? Stability problems don't suddenly happen. They are often the result of a long, slow grind to the point where stability (or performance) is unacceptable. Something happens. No one is quite sure. Reboot a few things - and the problem seems solved! A couple of weeks later, the reboots are more frequent. After a couple of months, the reboots don't seem to have any lasting effect - and a support incident gets opened. Incident tracking captures these seemingly unrelated events and generates a larger perspective. You still may not know what to do but you can see where it is trending. You may start to look for other correlations. You end up with a much better history and timeline of when and how things started going bad - and that will be a big help once it becomes a support incident.
It also makes it easy to confirm the fix is helping, or not, as the frequency of incidents changes.
As to the nature of the instability, and the effort to re-mediate - these are turning out to be real systemic problems - so no easy fixes. Getting Incident tracking re-established for the APM solution - that's easy enough - but the damage has already been done.
How do you track the performance and capacity of your APM solution? When do you think it will "matter"?
I believe this is due to the overall misunderstanding about APM being 'just another monitoring tool'. Some folks think APM looks like many other tools - just with better capabilities. And they treat it just the same - a tool for operations to use when there is a problem; and back on the shelf when things are quiet. This ignores the 24x7 reality of the solution. We already know that this presumption leaves significant gaps in managing the capacity of the metrics storage component.
A small APM initiative can go for years before they run out of capacity and no one is managing this until the solution becomes unstable - and then they realize they have limited understanding. A growing APM initiative will run into this problem more quickly, depending on the pace of their successive deployments. But the pace of deployment is not as significant as much as the absence or presence of incident tracking for the APM solution.
In the absence of incident tracking, the client continues on blindly experiencing instability of the monitoring solution and then will escalate to vendor support, initialling labeling everything as a 'product defect'. The vendor support will then attempt to confirm the 'defect' but after finding nothing wrong (no known incidents, no prior history of instability), you end up with a bit of an impasse: no defect, yet no resolution because the problem is the configuration, not the monitoring software.
Why then is 'incident tracking' so magical? Stability problems don't suddenly happen. They are often the result of a long, slow grind to the point where stability (or performance) is unacceptable. Something happens. No one is quite sure. Reboot a few things - and the problem seems solved! A couple of weeks later, the reboots are more frequent. After a couple of months, the reboots don't seem to have any lasting effect - and a support incident gets opened. Incident tracking captures these seemingly unrelated events and generates a larger perspective. You still may not know what to do but you can see where it is trending. You may start to look for other correlations. You end up with a much better history and timeline of when and how things started going bad - and that will be a big help once it becomes a support incident.
It also makes it easy to confirm the fix is helping, or not, as the frequency of incidents changes.
As to the nature of the instability, and the effort to re-mediate - these are turning out to be real systemic problems - so no easy fixes. Getting Incident tracking re-established for the APM solution - that's easy enough - but the damage has already been done.
How do you track the performance and capacity of your APM solution? When do you think it will "matter"?
Friday, December 17, 2010
Thursday, December 9, 2010
Some light at the end of the tunnel
Finished all writing and editing last week. This week brings the galley proofs.
Monday, May 3, 2010
On assignment on Australia
Now two weeks in to a four week assessment project, I think I'm getting a second round of jet lag.
The best thing about this trip is that, with the time difference, I can't get involved with anything else - so I can devote a lot of time to writing. There is 1 hr in the AM and 1 hr at night where there is enough overlap for a call - like that movie "LadyHawke" with Rutger Hauer. I finally got some chapters up to the publisher site, after spending too much time reworking and reorganizing content.
I've also figured out Australian rules football and the multiple rugby organizations. Knowledge that will probably help solve Climate Change...
The worst thing about the trip - I can only devote time to writing! And I'm writing the section about assessments... which I also doing in the service engagement... I need a vacation.
Note to self - full-time writing is not a career choice!
The best thing about this trip is that, with the time difference, I can't get involved with anything else - so I can devote a lot of time to writing. There is 1 hr in the AM and 1 hr at night where there is enough overlap for a call - like that movie "LadyHawke" with Rutger Hauer. I finally got some chapters up to the publisher site, after spending too much time reworking and reorganizing content.
I've also figured out Australian rules football and the multiple rugby organizations. Knowledge that will probably help solve Climate Change...
The worst thing about the trip - I can only devote time to writing! And I'm writing the section about assessments... which I also doing in the service engagement... I need a vacation.
Note to self - full-time writing is not a career choice!
Friday, March 19, 2010
Contract signed - now get back to writing!
Well, after many months of little progress the proposal has become project and we are with a signed contract with Apress to produce the book. I think we can get the whole project wrapped for the fall - maybe September. I still need to finalize the editorial schedule, and sign up some more internal reviewers.
Wednesday, December 16, 2009
Oblicore - Service Level Management
Notes from the public (registration required)online seminar 16 December 2009, 11:00 AM EST:
Deming (father of quality control) - Plan, Do, Check (measure), Act
SLA can mean many things - ITIL, ERP, CRM, COBiT - folks think they are talking about the same things but meaning varies, depending on where they are in the organization.
"When you have no destination - Every road leads nowhere". You need to align business needs with targets.
- An SLA is a contract between a provider and a customer. Service Level Agreement, Operational Level Agreement, Underpinning Contract. They all document targets and specifies responsibilities of the parties.
- SLM -- ITIL Service Level Management. Tells you what to do (but necessarily how to do it)
- Define, document, agree, monitor, measure, report and review the level of IT services provided.
Most companies manage SLAs manually. Each contract negotiated separately. Labor intensive data collection. Reporting is reactive.
- Establish - Implement - Manage - Review:
Service Improvement Plan: Define Strategy, Planning, Implementation, Catalog Services, Draft, Negotiate, Review SLAs SCs and OLAs, Agree, measure and monitor, Report, Review SLM process, Review SLAs SCs and OLAs (back to start of loop).
( lots about roles for the above steps...)
- Standardizing Contract Creation and Revision - Data collection and metrics need to be uniform.
-- most SLAs require multiple data sources. Need to aggregate the data, apply exceptions (TOD, external events), correlations with other incidents.
(more on types of pain when a robust process for SLA management is absent)
Oblicore does the last mile to drive an automated ITIL approach, via a double closed-loop process. (Geeze - this is buzzword heavy!).
(demo of Oblicore Guarantee) UI is browser-based, multiple tabs. Basically workflow management with a federated view of multiple metrics, over a various web-based forms. following the ITIL model. Generates paper! But somebody has to sign this stuff - so, way better than excel... They use adapters to bring in various metrics. Looks like an ETL transform (Table/Field assignment). Don't know how real-time this could be.
Deming (father of quality control) - Plan, Do, Check (measure), Act
SLA can mean many things - ITIL, ERP, CRM, COBiT - folks think they are talking about the same things but meaning varies, depending on where they are in the organization.
"When you have no destination - Every road leads nowhere". You need to align business needs with targets.
- An SLA is a contract between a provider and a customer. Service Level Agreement, Operational Level Agreement, Underpinning Contract. They all document targets and specifies responsibilities of the parties.
- SLM -- ITIL Service Level Management. Tells you what to do (but necessarily how to do it)
- Define, document, agree, monitor, measure, report and review the level of IT services provided.
Most companies manage SLAs manually. Each contract negotiated separately. Labor intensive data collection. Reporting is reactive.
- Establish - Implement - Manage - Review:
Service Improvement Plan: Define Strategy, Planning, Implementation, Catalog Services, Draft, Negotiate, Review SLAs SCs and OLAs, Agree, measure and monitor, Report, Review SLM process, Review SLAs SCs and OLAs (back to start of loop).
( lots about roles for the above steps...)
- Standardizing Contract Creation and Revision - Data collection and metrics need to be uniform.
-- most SLAs require multiple data sources. Need to aggregate the data, apply exceptions (TOD, external events), correlations with other incidents.
(more on types of pain when a robust process for SLA management is absent)
Oblicore does the last mile to drive an automated ITIL approach, via a double closed-loop process. (Geeze - this is buzzword heavy!).
(demo of Oblicore Guarantee) UI is browser-based, multiple tabs. Basically workflow management with a federated view of multiple metrics, over a various web-based forms. following the ITIL model. Generates paper! But somebody has to sign this stuff - so, way better than excel... They use adapters to bring in various metrics. Looks like an ETL transform (Table/Field assignment). Don't know how real-time this could be.
Friday, December 4, 2009
Storm Clouds are Clearing
Well the word is that we are refitted with an appropriate executive sponsor and a status call is scheduled for Monday 7 December. I'm still expecting a few bumps in the road.
Review - Symantec I3 - A Performance Management Methodology
[Clearing out the Drafts folder (Dec 4)]
http://silos-connect.com/solutions/i3methodologybook.pdf
This has a 2005 copyright - so I'm not expecting to get anything too useful here. But it does show up when you are searching on "APM Methodologies".
Introduction - Poor app performance translate into poor end-user experience.
Chapter 1 - Defines incidents as "performance red zones" - where the system fails to meet performance goals, and alludes that these may also be used to define SLAs. They break down incidents into four major types, and I've added some more conventional definition of what they had in mind:
Chapter 2 - provides the substance for the core activities described in the summary from chapter 1.
Reactive Management IS jumping to action when an incident occurs. But the call to action comes from an alert and the problem with alerts is that they very often don't bubble up to a Trouble Management console until some 15-30 minutes after the incident started. So somehow being closer to the keyboard doesn't really add any value - you are still the last to know. When your helpdesk is a more timely indicator of application health, than your monitoring solution - this is when you know you have a major visibility deficit.
Likewise, defining proactive as more frequent monitoring, in order to close that 15-30 minute gap is also a disservice. The operations team is simply not in a position to be constanting reviewing the key metrics for hundreds of applications under their watch. They rely on alerts to filter out the application noise (all those squiggly traces) and indicate which app is having problems. It might surprise you to also realize that operations has exactly two responses to an alert: 1. Restart the app, or 2. open a bridge call, and restart the app. That's 95% of modern IT today and that's the real problem.
I define proactive management as preventing the problem from getting deployed in the first place - not simply responding to an alert more quickly. The writer assigns this instead to preventative management, so maybe it is only a semantic (funny, no?) difference. But if they mean the periodic healthchecks as something that is occurring in operations, then this is a fantasy. Reverse-engineering the normal performance characteristics from an operational application is a massive task. Remember, we are not talking about the home web server - we are talking about a commercial, enterprise environment with hundreds of applications. That's really not the role for operations to undertake and decide. In order to make the problem tractable, the app owners have to pass judgment on what is normal, and what is abnormal for their app. In reality, they are the only ones who have a chance at understanding.
Regards the hybrid approach, that's something useful but only if we are feeding operational information back to Dev and Qa, in order to improve the initial testing of the app. And feeding forward results from QA, Stress and Performance, that can be used to set initial alert thresholds and maybe some info as to what known problems might actually look like, hopefully in terms of run documentations and/or dashboards.
The dashboard metaphor is key because no one has any time to figure out where the logs are for a given app. You need a mechanism to present details about the app, along with a list of who to call and what some of the resources involved might be.
Chapter 3 - Symantic Methodology - Process and Tools. Oops - I prefer People, Process and Product (tools). Must be some Borg over there at Symantec... And they claim the effectiveness of their "proven process" - but they haven't actually defined it yet... and then we are on to the products Insight, Indepth and Inform. And that's all for that section.
Performance Management Stages include detect (symptoms), find (source), focus (root-cause), improve (follow steps to improve) and verify (verify). I call that Triage and Remediation. My first step is not detect but refute! I prefer a less adversarial approach by first intoning "how do we know we have a problem?" It sometimes makes for some uncomfortable moments of silence.
Think about it. Someone notifies you that there is an urgent issue. How do you know that they are accurate? What do you look at to confirm that the issue is actually an incident? This is an important bit of process because the stages of explanation can become something consistent, like a practiced drill that can be executed when stress is high and patience is short. When everybody is used to it, it is actually quicker to review what's right and what is apparently wrong. And you need a steady soul to initiate this and keep it on track - but that's what effective triage is really about - deliberate, conclusive steps until something is found out of place. Not as much fun as running around with your hair on fire but a whole lot more effective and predictable.
One of the big complaints I hear from operations is that alerts are dropping in all the time. It's not just too many alerts and their frequency. It often means too many alerts for which nothing could be found indicating a problem. False-positives. Not actionable. Ultimately these are due to defects in visibility - and something that no amount of pressure and screaming can resolve. And there is an important point skipped over - the nature of alerts. Most alerts are that a system is unavailable - it has gone down or is unreachable. The supposed "pro-active" alerts - these are different because they have thresholds defined. It's not unfair to suggest that alerting on a threshold is more visibility than a "gone-down" alert - but it certainly isn;t "proactive" - it's using a threshold to define the alert. Duh! But how do you arrive at the threshold? What metrics do you select?
As the writer revisits the different management style, they point out that the mechanisms to realize the process improvements are part of the product (tools). Well, that's cool but what I really like to focus on are the product-neutral processes - the things I should be doing no matter whose tool set is in play. Sure, it's nice that you have an automated mechanism to periodically do performance reviews. What do you look at? What metrics are important? What changes are significant? How does this relate to what the operations team is needed to be more effective.
Process is something that is easy to wave around. Automated processes sound even better. What the process is and how it relates to the current organization and capabilities - I don't think the Symantec tool has any concept of that. And that is the gap that limits adoption and, ultimately, the utility of the tool.
In summary, the Symantec methodology is detect, find, focus, improve, and verify. The different products (tools) implement and automate these processes. I guess that could be anything.
Section 2 - devotes a chapter (chapters 4-8) to each step of the methodology. Performance Reviews are the mechanism: the periodic health check. This requires (4) types of reports: Top-N (response times), Trends and Exceptions. Apparently, there are only (3)! I guess the invocations Top-N report is implied. The trends are just a historical view of a metric - no magic here! The exceptions are an historical view of exceptions and errors - also not magical. There should be a status or overview fo the environment - again, no magic needed.
As the author moves back to Reactive Management, they introduce the concept of resource monitoring - databases, messaging, etc. These need specialized views (or tools), and no surprise here but can't we be "proactive" for resource management as well?
http://silos-connect.com/solutions/i3methodologybook.pdf
This has a 2005 copyright - so I'm not expecting to get anything too useful here. But it does show up when you are searching on "APM Methodologies".
Introduction - Poor app performance translate into poor end-user experience.
Chapter 1 - Defines incidents as "performance red zones" - where the system fails to meet performance goals, and alludes that these may also be used to define SLAs. They break down incidents into four major types, and I've added some more conventional definition of what they had in mind:
- Design and Deployment Problems (Design, Code, Build, Test, Package, Deploy)
- System Change Problems (Configuration mistakes due to tuning or attempted fix)
- Environment Change Problems (unexpected load or usage patterns)
- Usage Problems (Prioritized Access to Shared IT Resources)
"Performance management comprises three core activities, reactive, proactive and preventative. Symantic i3 provides the techniques and the tools that are necessary to carry out these activities using a structured, methodical, and holistic approach."As much as the core activities are accurate, there is nothing in chapter 1 that supports that summary.
Chapter 2 - provides the substance for the core activities described in the summary from chapter 1.
- Reactive Management - maximize alertness and minimize time-to-detection of problems, equip staff with the right tools to minimize time-to-resolution. I guess lot's of caffine is one of the tools!
- Proactive Management - "close monitoring" will often detect that problems are likely to occur. I guess you need to be in the same room as the monitoring...
- Preventative Management - Minimize performance problems in the first place by employing periodic health checks and resolve problems before they get out of hand.
- Hybrid Approach - Combine all (3) management styles, with feed-forward and feed-back, leading to an overall decline in firefighting situations.
Reactive Management IS jumping to action when an incident occurs. But the call to action comes from an alert and the problem with alerts is that they very often don't bubble up to a Trouble Management console until some 15-30 minutes after the incident started. So somehow being closer to the keyboard doesn't really add any value - you are still the last to know. When your helpdesk is a more timely indicator of application health, than your monitoring solution - this is when you know you have a major visibility deficit.
Likewise, defining proactive as more frequent monitoring, in order to close that 15-30 minute gap is also a disservice. The operations team is simply not in a position to be constanting reviewing the key metrics for hundreds of applications under their watch. They rely on alerts to filter out the application noise (all those squiggly traces) and indicate which app is having problems. It might surprise you to also realize that operations has exactly two responses to an alert: 1. Restart the app, or 2. open a bridge call, and restart the app. That's 95% of modern IT today and that's the real problem.
I define proactive management as preventing the problem from getting deployed in the first place - not simply responding to an alert more quickly. The writer assigns this instead to preventative management, so maybe it is only a semantic (funny, no?) difference. But if they mean the periodic healthchecks as something that is occurring in operations, then this is a fantasy. Reverse-engineering the normal performance characteristics from an operational application is a massive task. Remember, we are not talking about the home web server - we are talking about a commercial, enterprise environment with hundreds of applications. That's really not the role for operations to undertake and decide. In order to make the problem tractable, the app owners have to pass judgment on what is normal, and what is abnormal for their app. In reality, they are the only ones who have a chance at understanding.
Regards the hybrid approach, that's something useful but only if we are feeding operational information back to Dev and Qa, in order to improve the initial testing of the app. And feeding forward results from QA, Stress and Performance, that can be used to set initial alert thresholds and maybe some info as to what known problems might actually look like, hopefully in terms of run documentations and/or dashboards.
The dashboard metaphor is key because no one has any time to figure out where the logs are for a given app. You need a mechanism to present details about the app, along with a list of who to call and what some of the resources involved might be.
Chapter 3 - Symantic Methodology - Process and Tools. Oops - I prefer People, Process and Product (tools). Must be some Borg over there at Symantec... And they claim the effectiveness of their "proven process" - but they haven't actually defined it yet... and then we are on to the products Insight, Indepth and Inform. And that's all for that section.
Performance Management Stages include detect (symptoms), find (source), focus (root-cause), improve (follow steps to improve) and verify (verify). I call that Triage and Remediation. My first step is not detect but refute! I prefer a less adversarial approach by first intoning "how do we know we have a problem?" It sometimes makes for some uncomfortable moments of silence.
Think about it. Someone notifies you that there is an urgent issue. How do you know that they are accurate? What do you look at to confirm that the issue is actually an incident? This is an important bit of process because the stages of explanation can become something consistent, like a practiced drill that can be executed when stress is high and patience is short. When everybody is used to it, it is actually quicker to review what's right and what is apparently wrong. And you need a steady soul to initiate this and keep it on track - but that's what effective triage is really about - deliberate, conclusive steps until something is found out of place. Not as much fun as running around with your hair on fire but a whole lot more effective and predictable.
One of the big complaints I hear from operations is that alerts are dropping in all the time. It's not just too many alerts and their frequency. It often means too many alerts for which nothing could be found indicating a problem. False-positives. Not actionable. Ultimately these are due to defects in visibility - and something that no amount of pressure and screaming can resolve. And there is an important point skipped over - the nature of alerts. Most alerts are that a system is unavailable - it has gone down or is unreachable. The supposed "pro-active" alerts - these are different because they have thresholds defined. It's not unfair to suggest that alerting on a threshold is more visibility than a "gone-down" alert - but it certainly isn;t "proactive" - it's using a threshold to define the alert. Duh! But how do you arrive at the threshold? What metrics do you select?
As the writer revisits the different management style, they point out that the mechanisms to realize the process improvements are part of the product (tools). Well, that's cool but what I really like to focus on are the product-neutral processes - the things I should be doing no matter whose tool set is in play. Sure, it's nice that you have an automated mechanism to periodically do performance reviews. What do you look at? What metrics are important? What changes are significant? How does this relate to what the operations team is needed to be more effective.
Process is something that is easy to wave around. Automated processes sound even better. What the process is and how it relates to the current organization and capabilities - I don't think the Symantec tool has any concept of that. And that is the gap that limits adoption and, ultimately, the utility of the tool.
In summary, the Symantec methodology is detect, find, focus, improve, and verify. The different products (tools) implement and automate these processes. I guess that could be anything.
Section 2 - devotes a chapter (chapters 4-8) to each step of the methodology. Performance Reviews are the mechanism: the periodic health check. This requires (4) types of reports: Top-N (response times), Trends and Exceptions. Apparently, there are only (3)! I guess the invocations Top-N report is implied. The trends are just a historical view of a metric - no magic here! The exceptions are an historical view of exceptions and errors - also not magical. There should be a status or overview fo the environment - again, no magic needed.
As the author moves back to Reactive Management, they introduce the concept of resource monitoring - databases, messaging, etc. These need specialized views (or tools), and no surprise here but can't we be "proactive" for resource management as well?
Monday, November 30, 2009
The Essence of Triage
[clearing out the drafts folder - Nov 30]
When interpreting performance data for an incident the question often arises as to what should we look at first. For my APM practice I always focus on "what changed" and this is easily assessed by comparing with the performance baseline or signature for the affected application.
But for folks new to APM, and often very much focused on jumping in to look at individual metrics, you can easily get confused but so many metrics will be suspicious. There are some attributes of the application; response times, invocation rates, garbage collection, CPU, that will be out of normal. And folks will bias their recommendations as to which avenue to explore based on the experience they have with a particular set of metrics.
My approach to this is pretty basic: go for the "loudest" characteristic first. Much like "the squeeky wheel gets the oil" - the "loudest" metric is where you should start you investigation, moving downhill. More importantly, you need to stop looking once you have found a defect or deviation from your expectations, and get it fixed and then look again for the "loudest' metric.
This is important because the application will re-balance after the first problem is addressed, and you will get a new hierarchy of "loud" metrics to consider.
For example, let's assert a real-world scenario where there are two servlets, each of which is accessing a separate database. Use Case A accesses Servlet A, which is accessing an RDBMS and has a query response time of 4 seconds. Servlet A has a response time of 5 seconds. Use Case B has a response time of 2 seconds, accessing Servlet B and makes a query via messaging middleware to a mainframe HFS, which takes 1 second. Which of these is the loudest problem?
If you feel that a servlet response time of 5 seconds is a pretty good clue. You would be wrong. Sure, everyone should know that a servlet response time should be on the order of 1-3 seconds. And certainly being able to compare this performance to an established baseline would confirm it.
Instead, we will limit our consideration to the Use Case which has actually has users complaining, which in this case is Use case B.
"Wait a minute! You didn't tell us which use case had users complaining!"
Right. And neither will your users (real-world scenario)! What I'm driving at is that you can't know where to look until you know what has changed. And you can't know what has changed unless you have a normal baseline with which to compare. It's always nice when you have an alert or user complaint to help point you in the right direction but that can be unreliable as well.
For all this I prefer what I call the "Dr. House" model. Dr. House is a TV show what House draws out the root cause for troublesome medical cases he and his team are involved with. I think it's a great model for triage and diagnosis of application performance. One of the axioms of House's interactions with patients is that "everybody lies", when they present their medical history.
This is how I approach triage - everybody lies (or in corporate-neutral language: inadvertently withholds key information). So I base much of my conclusions as to how to proceed and what to look for, based on what I can expose by comparing the current activity, to 'normal' - or whatever baseline I can develop.
When interpreting performance data for an incident the question often arises as to what should we look at first. For my APM practice I always focus on "what changed" and this is easily assessed by comparing with the performance baseline or signature for the affected application.
But for folks new to APM, and often very much focused on jumping in to look at individual metrics, you can easily get confused but so many metrics will be suspicious. There are some attributes of the application; response times, invocation rates, garbage collection, CPU, that will be out of normal. And folks will bias their recommendations as to which avenue to explore based on the experience they have with a particular set of metrics.
My approach to this is pretty basic: go for the "loudest" characteristic first. Much like "the squeeky wheel gets the oil" - the "loudest" metric is where you should start you investigation, moving downhill. More importantly, you need to stop looking once you have found a defect or deviation from your expectations, and get it fixed and then look again for the "loudest' metric.
This is important because the application will re-balance after the first problem is addressed, and you will get a new hierarchy of "loud" metrics to consider.
For example, let's assert a real-world scenario where there are two servlets, each of which is accessing a separate database. Use Case A accesses Servlet A, which is accessing an RDBMS and has a query response time of 4 seconds. Servlet A has a response time of 5 seconds. Use Case B has a response time of 2 seconds, accessing Servlet B and makes a query via messaging middleware to a mainframe HFS, which takes 1 second. Which of these is the loudest problem?
If you feel that a servlet response time of 5 seconds is a pretty good clue. You would be wrong. Sure, everyone should know that a servlet response time should be on the order of 1-3 seconds. And certainly being able to compare this performance to an established baseline would confirm it.
Instead, we will limit our consideration to the Use Case which has actually has users complaining, which in this case is Use case B.
"Wait a minute! You didn't tell us which use case had users complaining!"
Right. And neither will your users (real-world scenario)! What I'm driving at is that you can't know where to look until you know what has changed. And you can't know what has changed unless you have a normal baseline with which to compare. It's always nice when you have an alert or user complaint to help point you in the right direction but that can be unreliable as well.
For all this I prefer what I call the "Dr. House" model. Dr. House is a TV show what House draws out the root cause for troublesome medical cases he and his team are involved with. I think it's a great model for triage and diagnosis of application performance. One of the axioms of House's interactions with patients is that "everybody lies", when they present their medical history.
This is how I approach triage - everybody lies (or in corporate-neutral language: inadvertently withholds key information). So I base much of my conclusions as to how to proceed and what to look for, based on what I can expose by comparing the current activity, to 'normal' - or whatever baseline I can develop.
Wednesday, November 25, 2009
Yikes - project crashing!
My executive sponsor has bailed on the project. No details yet as to why. I am bummed. Sure, just need to sign up a fresh one but time marches on...
Tuesday, November 24, 2009
What You Should Know... part 2
Well, this part was less annoying.
All in all, what can I say about this author? He has 30 books published. I have a book proposal. So I am crap.
But. He writes about APM. I "do" APM. I help clients realize APM. I "talk the talk" AND "walk the walk".
The author is "two-thousand and LATE"! ;-)
Now, if I only had a book...
- Application lifecycle - includes development. Tru-dat my brother! Too bad the authors example is of a code-profiler. And yes, most APM-savvy folks do not include development as part of APMonitoring. But if you really want to improve app performance (and the end-user experience) - you need cooperation from development.
- The SOA flag - This is not so bad. He inserts that ASM acronym again. But otherwise, this is accurate and helpful. You know, if you really think that APM is overloaded (which it is) - how about SPM - Service Performance Management. Everybody knows "ASM" means an assembler directive anyway!
All in all, what can I say about this author? He has 30 books published. I have a book proposal. So I am crap.
But. He writes about APM. I "do" APM. I help clients realize APM. I "talk the talk" AND "walk the walk".
The author is "two-thousand and LATE"! ;-)
Now, if I only had a book...
What You Should Know About Application Performance management
This is from RealTime Nexus - The Digital Library for IT Professionals. Do your own search - I will not spoil my page with a link to it! Let me opine on the salient points:
In assessing the utility of an APM initiative, the focus is always on the high-value transactions - end-user-related or not. Then when you know what really matters, you select the appropriate technology. That way, you do end up trying to use end-user monitoring on a CICS transaction, nor using byte-code instrumentation to monitor a print server. ;-)
"Application-level resource monitors, if any exist. These have to be specifically created and exposed by the application author."
I guess they never heard of BCI (Byte Code Instrumentation) - which does this automatically. And have impuned JMX and PMI technologies - which do the right job for configuration information of the application server - which is what I'm hoping the author really meant. JMX and PMI require the developer to code for their use. Always was and always will - and an expensive proposition at development. But BCI automatically determines what is interesting to monitor, much more effectively that JMX or PMI - and at runtime (aka - late-binding). But if the data is already there - we take direct advantage of it.
- Makes a distinction for APMonitoring (APM == Monitoring) and somehow reserves the assignment of thresholds to Health Monitoring, as different from Performance Monitoring, even suggesting "AHM", as a new acronym... but then realizing that APM is already well accepted. I think it would be more accurate to acknowledge that APM tools don't do the "Health Monitoring" - OOTB, but I would submit that APM processes would address this - especially as I already use a "HealthCheck" process as part of the APManagement best practices.
- Asserts that "end-user performance" is the primary metric for APM and acknowledges that the "other metric may be involved... to provide troubleshooting details. " This is too much of a plug for a specific vendor solution. Sure end-user experience is what performance management is all about but it is a little naive to assert that it is the main focus. There is a huge benefit in using APM across the application lifecycle, and especially before the end-user experience can even be measured (development, testing). Not to mention the significant number of applications that do not have even have a user front-end to measure!
In assessing the utility of an APM initiative, the focus is always on the high-value transactions - end-user-related or not. Then when you know what really matters, you select the appropriate technology. That way, you do end up trying to use end-user monitoring on a CICS transaction, nor using byte-code instrumentation to monitor a print server. ;-)
- How does APM work -- I was nestling in for a good read here - but was disappointed. Regarding the types of information that an APM tool can take advantage, the author describes the following:
"Application-level resource monitors, if any exist. These have to be specifically created and exposed by the application author."
I guess they never heard of BCI (Byte Code Instrumentation) - which does this automatically. And have impuned JMX and PMI technologies - which do the right job for configuration information of the application server - which is what I'm hoping the author really meant. JMX and PMI require the developer to code for their use. Always was and always will - and an expensive proposition at development. But BCI automatically determines what is interesting to monitor, much more effectively that JMX or PMI - and at runtime (aka - late-binding). But if the data is already there - we take direct advantage of it.
- Downsides of APM -- This is annoying because it is a grain of truth buried in a FoxNews-spun positioning. Sure, packaged apps are hard to monitor but this is because they are closed and usually provide their own monitoring tools. They may not be best of breed - but it's a start. Ratified packaged vendors will actually embeed APM technologies within their offerings, and some require a vendor-specific extension - for sure, the industry is lagging a bit here - and that's always a problem with proprietary solutions. It is not an APM problem.
- Application Service Management (ASM) -- this is the point that set me off and motivated me to opine this detailed review. The author creates a new acronym - cool. I do it all the time. No crime here. But "ASM takes a more business-oriented view." - yikes! ASM focuses on the operating system and platform metrics... I though that's NOT the business view. And then the author acknowledges that the APM/ASM differences are "semantic" - and you should never get caught up in that. Frankly, this is the same tact that creationists take when they assert that "... evolution is a theory that is still being considered...". Dude!!!
Wednesday, November 18, 2009
Legal approved the project, with caveats
With little fanfare the first major hurdle has been surpassed. The caveats were as expected: no discussion of product technology, no vendor or client names, and no vendor bias - what we call vendor-neutral. The last point has been the real 'chestnut' for the marketing folks with past projects. They want to control the product message and spin things in their direction. I've always been of the mind that the only real solution was to stay vendor and technology neutral. I've been doing it this way for years.
As I've been sticking to "vendor-neutral" as the dominant theme for the APM practice, over the last (4) years and it has clearly been the "path least traveled". If there is nothing to leverage, the marketing arm has generally been indifferent to the program but sometimes it would be a little more pointed. The name of the program, for instance, went through a number of iterations, each about (9) months apart, and ending with a pronouncement and re-branding of all the materials - and then silence. And when we did get some marketing support for a brochure or web link, you really had to hunt it down.
I had last year an MBO to develop sales training materials, get it recorded and up on our internal training site. Everything was under "severe urgency" and when it was finished, no notice of the update was allowed and no edit to make it part of the presales learning plan - which is really the only way folks will spend a couple hours doing training. Sure, it was a month late but I had some family problems. So here it is a full year later and folks are still surprised that they can find it.
So just three days later, after the blessing by legal, I'm now finding that the Project PM is questioning the very premise of re-purposing the root documents for the practice, into book form, and that call will not happen until December 4th. This will suck!
Call me paranoid but I think the old product-centric nemesis is back again! - preserve the status quo, it's not our business anyway, it will only help the competition survive another year...
I did a best practice paper a couple years back focusing on memory management and going way more than the product doc to really show folks how to use it correctly and, more importantly, how to back out of problems when it was used incorrectly. A year as a proposal and a cherry of a project (from a technical perspective), I pick up the project (original lead back off the project) and finished in 3 months. Multiple reviewers and probably my best work at the time. Then it went to marketing for approval to publish and sat there for (9) months. Then it was released, with no edits at all or commentary. My inside sources said the issue was that marketing could not accept that someone could misconfigure the product. Sounds plausible - but folks blown stuff up all the time, for any software system. It seems cruel that you would not show them how to back out gracefully - but that's just my opinion. Anyway, I never sent any more material up to marketing. And now that decision is back to bite me - and with 4 years of constant writing, without any marketing oversight, will they actually review it? Shoot the project? Or throw up their hands and yield to the marketplace?
My original response to management, back in August when the book was 'commissioned', was that marketing would wake up and block a book so why not look at just dropping the fame and glory and go with the Google Knol? We publish what we want (with blessing by legal), the world is made a better place - and I can go on with the next big thing. I still think that is the best route, especially as my optimism wanes.
Maybe it will be better after some turkey?
As I've been sticking to "vendor-neutral" as the dominant theme for the APM practice, over the last (4) years and it has clearly been the "path least traveled". If there is nothing to leverage, the marketing arm has generally been indifferent to the program but sometimes it would be a little more pointed. The name of the program, for instance, went through a number of iterations, each about (9) months apart, and ending with a pronouncement and re-branding of all the materials - and then silence. And when we did get some marketing support for a brochure or web link, you really had to hunt it down.
I had last year an MBO to develop sales training materials, get it recorded and up on our internal training site. Everything was under "severe urgency" and when it was finished, no notice of the update was allowed and no edit to make it part of the presales learning plan - which is really the only way folks will spend a couple hours doing training. Sure, it was a month late but I had some family problems. So here it is a full year later and folks are still surprised that they can find it.
So just three days later, after the blessing by legal, I'm now finding that the Project PM is questioning the very premise of re-purposing the root documents for the practice, into book form, and that call will not happen until December 4th. This will suck!
Call me paranoid but I think the old product-centric nemesis is back again! - preserve the status quo, it's not our business anyway, it will only help the competition survive another year...
I did a best practice paper a couple years back focusing on memory management and going way more than the product doc to really show folks how to use it correctly and, more importantly, how to back out of problems when it was used incorrectly. A year as a proposal and a cherry of a project (from a technical perspective), I pick up the project (original lead back off the project) and finished in 3 months. Multiple reviewers and probably my best work at the time. Then it went to marketing for approval to publish and sat there for (9) months. Then it was released, with no edits at all or commentary. My inside sources said the issue was that marketing could not accept that someone could misconfigure the product. Sounds plausible - but folks blown stuff up all the time, for any software system. It seems cruel that you would not show them how to back out gracefully - but that's just my opinion. Anyway, I never sent any more material up to marketing. And now that decision is back to bite me - and with 4 years of constant writing, without any marketing oversight, will they actually review it? Shoot the project? Or throw up their hands and yield to the marketplace?
My original response to management, back in August when the book was 'commissioned', was that marketing would wake up and block a book so why not look at just dropping the fame and glory and go with the Google Knol? We publish what we want (with blessing by legal), the world is made a better place - and I can go on with the next big thing. I still think that is the best route, especially as my optimism wanes.
Maybe it will be better after some turkey?
Wednesday, October 28, 2009
Compuware - The Definitive Guide to APM
This is an web book from RealTime publishers nexus.realtimepublishers.com Currently, only half of the chapters are delivered, so here is what it is!
The Good: Emphasis on the process gap and organization maturity as the real barriers to APM success. It includes small vignettes of IT life at the start of each chapter, to highlight to IT situation and challenges. They even define the "M" of APM as Management, not Monitoring. They avoid the word "Dashboards" in favor of visualizations (cool: brings reports back to the table) and mention "application lifecycle" a few times, like "APM optimizes the application lifecycle". And, my favorite: "measuring performance is a proxy for understanding business performance".
The Bad: It purports to "take you through an entire implementation" but doesn't offer any depth. It is more of an extended whitepaper. It does cover the lifecycle of the motivation, decisions and implementation of an APM solution but as a conversation of what could be done. They only acknowledge stakeholders as SysAdmins, Developers and End-users (the guy using the browser)
The revelation for me was that they dredged up a Gartner Maturity Model from 2003 that had some interesting contrasts with the model we derived from our internal analysis of implementation failures and successes. Gartner identified management maturity as "Chaotic, Reactive, Proactive, Service and Value". Our model allowed for "Reactive, Directed, Proactive, Service-Driven and Value-Driven", with Reactive further divided as Reactive-Negotiated and Reactive-Alerting. I don't get too much access to Gartner stuff but here is how I represented Management Maturity:
This is from the first ICMM positioning around 2005. I don't recall why I decreased the size of each box, from left to right. I think I was trying to emphasize efficiency or proportion of IT organizations that might be found practicing at that level.
"Directed" is using APM metrics to influence the application lifecycle, focusing on QA practices.
"Service-driven" - everybody has this goal, we only tried to put teeth into it by associating it the definition of best practices.
I really liked this slide but it is not the emphasis we have today. We focus more on "visibility" because it provides a more "joining" context among the various tools that are available and their contributions, rather a focus on a particular technology and excluding all others. We also highlight the impact that visibility has on the existing processes in an organization and how this helps us assess their maturity and make meaningful recommendations for remediation.
Management Maturity and the processes that go along with it are the foundation of the Compuware APM message. That's not so bad. But then they fall into some unusal partitioning of APM in order to highlight their transaction monitoring technology. So I conclude that while they are saying the right things they regrettably recast APM as something that their technology delivers - and you only need to look at transactions.
Process re-engineering is actually pretty difficult and for well-established and reasonably successful, Reactive-Management organizations, this message is just way too hollow. These prospects know a lot about processes and may even know where they have gaps. What they need instead is a plan to help them evolve the organization, consistent with APM. What can they do now, and tomorrow, (and without purchasing new technology) that will get them on track to APM? What will derail their efforts? How will they know they have improved the situation? When will investment help accelerate their drive to APM?
The Good: Emphasis on the process gap and organization maturity as the real barriers to APM success. It includes small vignettes of IT life at the start of each chapter, to highlight to IT situation and challenges. They even define the "M" of APM as Management, not Monitoring. They avoid the word "Dashboards" in favor of visualizations (cool: brings reports back to the table) and mention "application lifecycle" a few times, like "APM optimizes the application lifecycle". And, my favorite: "measuring performance is a proxy for understanding business performance".
The Bad: It purports to "take you through an entire implementation" but doesn't offer any depth. It is more of an extended whitepaper. It does cover the lifecycle of the motivation, decisions and implementation of an APM solution but as a conversation of what could be done. They only acknowledge stakeholders as SysAdmins, Developers and End-users (the guy using the browser)
The revelation for me was that they dredged up a Gartner Maturity Model from 2003 that had some interesting contrasts with the model we derived from our internal analysis of implementation failures and successes. Gartner identified management maturity as "Chaotic, Reactive, Proactive, Service and Value". Our model allowed for "Reactive, Directed, Proactive, Service-Driven and Value-Driven", with Reactive further divided as Reactive-Negotiated and Reactive-Alerting. I don't get too much access to Gartner stuff but here is how I represented Management Maturity:
This is from the first ICMM positioning around 2005. I don't recall why I decreased the size of each box, from left to right. I think I was trying to emphasize efficiency or proportion of IT organizations that might be found practicing at that level."Directed" is using APM metrics to influence the application lifecycle, focusing on QA practices.
"Service-driven" - everybody has this goal, we only tried to put teeth into it by associating it the definition of best practices.
I really liked this slide but it is not the emphasis we have today. We focus more on "visibility" because it provides a more "joining" context among the various tools that are available and their contributions, rather a focus on a particular technology and excluding all others. We also highlight the impact that visibility has on the existing processes in an organization and how this helps us assess their maturity and make meaningful recommendations for remediation.
Management Maturity and the processes that go along with it are the foundation of the Compuware APM message. That's not so bad. But then they fall into some unusal partitioning of APM in order to highlight their transaction monitoring technology. So I conclude that while they are saying the right things they regrettably recast APM as something that their technology delivers - and you only need to look at transactions.
Process re-engineering is actually pretty difficult and for well-established and reasonably successful, Reactive-Management organizations, this message is just way too hollow. These prospects know a lot about processes and may even know where they have gaps. What they need instead is a plan to help them evolve the organization, consistent with APM. What can they do now, and tomorrow, (and without purchasing new technology) that will get them on track to APM? What will derail their efforts? How will they know they have improved the situation? When will investment help accelerate their drive to APM?
Subscribe to:
Posts (Atom)


