Seems I missed an upload during the final publishing push. While I sort that out with the publisher, here are the various docs that the book refers to, at Google Docs:
https://docs.google.com/leaf?id=0B9WYLZErvx39NzZlZjNjYTMtNjYyOS00NzcwLWIwMDQtZWYxNDE0ODYxNGE4&hl=en&authkey=CPjs-tEM
Saturday, March 12, 2011
Friday, February 18, 2011
Incident Tracking and APM Maturity
I've being doing more Architecture Assessments (post-deployment) than Planning Assessments lately and have noticed something troubling. Folks in general have mature incident and trouble management practices for the applications that they operate/manage. Yet they do not do any tracking for any incidents related to the APM solution. What's going on?
I believe this is due to the overall misunderstanding about APM being 'just another monitoring tool'. Some folks think APM looks like many other tools - just with better capabilities. And they treat it just the same - a tool for operations to use when there is a problem; and back on the shelf when things are quiet. This ignores the 24x7 reality of the solution. We already know that this presumption leaves significant gaps in managing the capacity of the metrics storage component.
A small APM initiative can go for years before they run out of capacity and no one is managing this until the solution becomes unstable - and then they realize they have limited understanding. A growing APM initiative will run into this problem more quickly, depending on the pace of their successive deployments. But the pace of deployment is not as significant as much as the absence or presence of incident tracking for the APM solution.
In the absence of incident tracking, the client continues on blindly experiencing instability of the monitoring solution and then will escalate to vendor support, initialling labeling everything as a 'product defect'. The vendor support will then attempt to confirm the 'defect' but after finding nothing wrong (no known incidents, no prior history of instability), you end up with a bit of an impasse: no defect, yet no resolution because the problem is the configuration, not the monitoring software.
Why then is 'incident tracking' so magical? Stability problems don't suddenly happen. They are often the result of a long, slow grind to the point where stability (or performance) is unacceptable. Something happens. No one is quite sure. Reboot a few things - and the problem seems solved! A couple of weeks later, the reboots are more frequent. After a couple of months, the reboots don't seem to have any lasting effect - and a support incident gets opened. Incident tracking captures these seemingly unrelated events and generates a larger perspective. You still may not know what to do but you can see where it is trending. You may start to look for other correlations. You end up with a much better history and timeline of when and how things started going bad - and that will be a big help once it becomes a support incident.
It also makes it easy to confirm the fix is helping, or not, as the frequency of incidents changes.
As to the nature of the instability, and the effort to re-mediate - these are turning out to be real systemic problems - so no easy fixes. Getting Incident tracking re-established for the APM solution - that's easy enough - but the damage has already been done.
How do you track the performance and capacity of your APM solution? When do you think it will "matter"?
I believe this is due to the overall misunderstanding about APM being 'just another monitoring tool'. Some folks think APM looks like many other tools - just with better capabilities. And they treat it just the same - a tool for operations to use when there is a problem; and back on the shelf when things are quiet. This ignores the 24x7 reality of the solution. We already know that this presumption leaves significant gaps in managing the capacity of the metrics storage component.
A small APM initiative can go for years before they run out of capacity and no one is managing this until the solution becomes unstable - and then they realize they have limited understanding. A growing APM initiative will run into this problem more quickly, depending on the pace of their successive deployments. But the pace of deployment is not as significant as much as the absence or presence of incident tracking for the APM solution.
In the absence of incident tracking, the client continues on blindly experiencing instability of the monitoring solution and then will escalate to vendor support, initialling labeling everything as a 'product defect'. The vendor support will then attempt to confirm the 'defect' but after finding nothing wrong (no known incidents, no prior history of instability), you end up with a bit of an impasse: no defect, yet no resolution because the problem is the configuration, not the monitoring software.
Why then is 'incident tracking' so magical? Stability problems don't suddenly happen. They are often the result of a long, slow grind to the point where stability (or performance) is unacceptable. Something happens. No one is quite sure. Reboot a few things - and the problem seems solved! A couple of weeks later, the reboots are more frequent. After a couple of months, the reboots don't seem to have any lasting effect - and a support incident gets opened. Incident tracking captures these seemingly unrelated events and generates a larger perspective. You still may not know what to do but you can see where it is trending. You may start to look for other correlations. You end up with a much better history and timeline of when and how things started going bad - and that will be a big help once it becomes a support incident.
It also makes it easy to confirm the fix is helping, or not, as the frequency of incidents changes.
As to the nature of the instability, and the effort to re-mediate - these are turning out to be real systemic problems - so no easy fixes. Getting Incident tracking re-established for the APM solution - that's easy enough - but the damage has already been done.
How do you track the performance and capacity of your APM solution? When do you think it will "matter"?
Friday, December 17, 2010
Thursday, December 9, 2010
Some light at the end of the tunnel
Finished all writing and editing last week. This week brings the galley proofs.
Monday, May 3, 2010
On assignment on Australia
Now two weeks in to a four week assessment project, I think I'm getting a second round of jet lag.
The best thing about this trip is that, with the time difference, I can't get involved with anything else - so I can devote a lot of time to writing. There is 1 hr in the AM and 1 hr at night where there is enough overlap for a call - like that movie "LadyHawke" with Rutger Hauer. I finally got some chapters up to the publisher site, after spending too much time reworking and reorganizing content.
I've also figured out Australian rules football and the multiple rugby organizations. Knowledge that will probably help solve Climate Change...
The worst thing about the trip - I can only devote time to writing! And I'm writing the section about assessments... which I also doing in the service engagement... I need a vacation.
Note to self - full-time writing is not a career choice!
The best thing about this trip is that, with the time difference, I can't get involved with anything else - so I can devote a lot of time to writing. There is 1 hr in the AM and 1 hr at night where there is enough overlap for a call - like that movie "LadyHawke" with Rutger Hauer. I finally got some chapters up to the publisher site, after spending too much time reworking and reorganizing content.
I've also figured out Australian rules football and the multiple rugby organizations. Knowledge that will probably help solve Climate Change...
The worst thing about the trip - I can only devote time to writing! And I'm writing the section about assessments... which I also doing in the service engagement... I need a vacation.
Note to self - full-time writing is not a career choice!
Friday, March 19, 2010
Contract signed - now get back to writing!
Well, after many months of little progress the proposal has become project and we are with a signed contract with Apress to produce the book. I think we can get the whole project wrapped for the fall - maybe September. I still need to finalize the editorial schedule, and sign up some more internal reviewers.
Wednesday, December 16, 2009
Oblicore - Service Level Management
Notes from the public (registration required)online seminar 16 December 2009, 11:00 AM EST:
Deming (father of quality control) - Plan, Do, Check (measure), Act
SLA can mean many things - ITIL, ERP, CRM, COBiT - folks think they are talking about the same things but meaning varies, depending on where they are in the organization.
"When you have no destination - Every road leads nowhere". You need to align business needs with targets.
- An SLA is a contract between a provider and a customer. Service Level Agreement, Operational Level Agreement, Underpinning Contract. They all document targets and specifies responsibilities of the parties.
- SLM -- ITIL Service Level Management. Tells you what to do (but necessarily how to do it)
- Define, document, agree, monitor, measure, report and review the level of IT services provided.
Most companies manage SLAs manually. Each contract negotiated separately. Labor intensive data collection. Reporting is reactive.
- Establish - Implement - Manage - Review:
Service Improvement Plan: Define Strategy, Planning, Implementation, Catalog Services, Draft, Negotiate, Review SLAs SCs and OLAs, Agree, measure and monitor, Report, Review SLM process, Review SLAs SCs and OLAs (back to start of loop).
( lots about roles for the above steps...)
- Standardizing Contract Creation and Revision - Data collection and metrics need to be uniform.
-- most SLAs require multiple data sources. Need to aggregate the data, apply exceptions (TOD, external events), correlations with other incidents.
(more on types of pain when a robust process for SLA management is absent)
Oblicore does the last mile to drive an automated ITIL approach, via a double closed-loop process. (Geeze - this is buzzword heavy!).
(demo of Oblicore Guarantee) UI is browser-based, multiple tabs. Basically workflow management with a federated view of multiple metrics, over a various web-based forms. following the ITIL model. Generates paper! But somebody has to sign this stuff - so, way better than excel... They use adapters to bring in various metrics. Looks like an ETL transform (Table/Field assignment). Don't know how real-time this could be.
Deming (father of quality control) - Plan, Do, Check (measure), Act
SLA can mean many things - ITIL, ERP, CRM, COBiT - folks think they are talking about the same things but meaning varies, depending on where they are in the organization.
"When you have no destination - Every road leads nowhere". You need to align business needs with targets.
- An SLA is a contract between a provider and a customer. Service Level Agreement, Operational Level Agreement, Underpinning Contract. They all document targets and specifies responsibilities of the parties.
- SLM -- ITIL Service Level Management. Tells you what to do (but necessarily how to do it)
- Define, document, agree, monitor, measure, report and review the level of IT services provided.
Most companies manage SLAs manually. Each contract negotiated separately. Labor intensive data collection. Reporting is reactive.
- Establish - Implement - Manage - Review:
Service Improvement Plan: Define Strategy, Planning, Implementation, Catalog Services, Draft, Negotiate, Review SLAs SCs and OLAs, Agree, measure and monitor, Report, Review SLM process, Review SLAs SCs and OLAs (back to start of loop).
( lots about roles for the above steps...)
- Standardizing Contract Creation and Revision - Data collection and metrics need to be uniform.
-- most SLAs require multiple data sources. Need to aggregate the data, apply exceptions (TOD, external events), correlations with other incidents.
(more on types of pain when a robust process for SLA management is absent)
Oblicore does the last mile to drive an automated ITIL approach, via a double closed-loop process. (Geeze - this is buzzword heavy!).
(demo of Oblicore Guarantee) UI is browser-based, multiple tabs. Basically workflow management with a federated view of multiple metrics, over a various web-based forms. following the ITIL model. Generates paper! But somebody has to sign this stuff - so, way better than excel... They use adapters to bring in various metrics. Looks like an ETL transform (Table/Field assignment). Don't know how real-time this could be.
Subscribe to:
Posts (Atom)