process training industry best practice problem management
agenda 1 2 3 4 5 6 7 our goal purpose of problem management general role description problem management flow best practices: problem detection best practices: problem prioritization best practices: problem investigation & diagnosis 8 9 best practices: problem resolution best practices: problem closure
1 our goal itil is the de facto standard for the whole it industry. however the itil release cycles vary from 5-10 years and do not reflect the speed of business changed today. our association focuses on cross-company collaboration to define best practices based on our design and operational experiences. our goal is to provide practical industry best practices beyond the itil standards. our chapter for today : problem management!
2 purpose of problem management itil quote: “the purpose of problem management is to manage the lifecycle of all problems from first identification through further investigation, documen- tation and eventual removal. problem management seeks to minimize the adverse impact of incidents and problems on the business that are caused by underlying errors within the it-infrastructure, and to proactively prevent recurrence of incidents related to these errors. in order to achieve this, problem management seeks to get to the root cause of incidents, documents and communicates known errors and initiates actions to improve or correct the situation.”
4 general role description service delivery manager (sdm) operations manager problem manager problem analyst main interface to customer allocates all customer requests into the internal organization. responsible for contract fulfill- ment regarding delivery and commercial aspects plans, controls and is respon- sible for the adequate provision of one or a number of services has operative responsibility for the whole lifecycle of the problems main contact for the sdm provides the technical inventory and service data for the part of the service ensures required approvals (root cause analysis results, resolution actions, etc..) leads the root cause analysis including detailed documen- tation and involvement of external parties ensures of all necessary resources and of all required information sources change manager change coordinator change approver change requester analyzes and reviews rfc and change planning coordinates changes and is responsible for change planning is responsible for safeguarding of complex and high risk changes approves changes; leads cab ensures mandatory approvals (e.g. presents the changes to man- datory change advisory boards ) is responsible that change implementation is compliant with approved change planning e.g. sdm, cabs, customer, etc. reviews the change planning approves or denies the change e.g. customer, product manager, project manager, sdm, etc. requests a change with a de- scription of the cause and the target as a rfc (request for change) change implementer incident manager lead incident manager (lim) manager on duty (mod) is responsible for activities as part of a change implementation ensures processing of incidents during their entire life cycle within service level or operational level agreements § overall responsibility for the execution of the inm process responsible to handle escalations steers as a project manager the incident solution is responsible for the document- ation and handover to problem management supports incident solution with dedicated customer or service line knowledge ensures that incident solution will be pushed forward 24x7
4 problem management flow based on itil service desk event management incident management proactive problem management supplier or contractor incident management implement workaround yes workaround needed? problem detection handover from major incident no raise known error record if required known error database change management rfc yes change needed? problem logging problem categorization problem prioritization supplier management cms problem investigation and diagnosis review changes if these caused the incident mi rca review & approval no problem resolution no resolved? yes problem closure major problem yes major problem review no end lessons learned review results service knowledge manage- ment system continual service improvement improvement actions communication
5 best practices: proactive problem management adapted problem flow service desk event management incident management proactive problem management supplier or contractor incident management implement workaround yes workaround needed? problem detection handover from major incident no raise known error record if required known error database change management rfc yes change needed? problem logging problem categorization problem prioritization supplier management cms problem investigation and diagnosis review changes if these caused the incident mi rca review & approval no problem resolution no resolved? yes problem closure major problem yes major problem review no end lessons learned review results service knowledge manage- ment system continual service improvement improvement actions communication
5 best practices: problem detection reactive problem management: analysis of recurring events and incidents recurring events & incidents can represent more than 50% of the whole incident amount. therefore it is important to identify similar incidents that might have the same root-cause, to find and remove the cause, so that the incidents will reoccur. steps to perform an analysis of recurring events and incidents: get a report with event & incident data as well as potential system health data from suppliers cluster incidents in homogeneous groups based on their description and cluster events & incidents according to their config items (e.g. by using pivot table) open and process problem tickets for each identified cluster important to identify smilar incident that migth have the same root-cause, find that cause and remove it, so config item 1 config item 2 config item 3 incident 1212 event 82123 event 82126 incident 8758 incident 2678 event 82129 config item 4 incident 8687 incident 2323 event 82155 event 82178 incident 8545 event 82123 event 82124 event 82125 event 82126 event 82127 event 82128 the multiple incidents can occur on more config items... (we can search for similarity of the symptoms e.g. via comparing the brief description of incidents) ...or on a single config item (the group of incidents may indicate some malfunction of the config item)
5 best practices: problem detection reactive problem management: handover from incident management 1/2 handover to problem management shall happen after the incident is solved or the service is stabilized with a workaround problem management is mandatory for all critical incidents (mi) and high incidents. we recommend to perform a „warm“ handover for at least major incidents into problem management. incident management owns the handover responsibility. this handover must include a time log (trigger of events) including the corresponding time zone for each entry. organization of a warm handover: a warm handover is taking place during an incident review call, which is necessary after a critical or high incident up on a special request. the review normally takes place after the last call, together with all technical and management key players of the related incident and is hosted by the manager on duty or lead incident manager which managed the incident. a warm handover is important to ensure that all important information will be transferred from incident management to problem management.
5 best practices: problem detection reactive problem management: handover from incident management 2/2 final incident report for warm handover: basis for the mi review is the „final incident report“, which needs to be shared to all participants. target is to review the complete incident history and is focused on: - what happened and when did it happen? - are the trigger of events correct and complete? - what led to the solution of the incident and at which exact time? - what are the points problem management has to focus on? - did we identify weak points during the major incident process? - which people are necessary for problem management? all topics shall be recorded in the “incident report”. define which time zone times are to be displayed and used in the report. recommendation is to use utc. major incident review has to be done during the office hours from monday to friday. example: incident occurs 08:21 08:59 incident recorded at service desk 10:45 major incident procedure initiated 11:45 first technician call 12:00 layer check initiated to identify the issue 14:57 first management call ….etc…
6 best practices: problem prioritization adapted problem flow service desk event management incident management proactive problem management supplier or contractor incident management implement workaround yes workaround needed? problem detection handover from major incident no raise known error record if required known error database change management rfc yes change needed? problem logging problem categorization problem prioritization supplier management cms problem investigation and diagnosis review changes if these caused the incident mi rca review & approval no problem resolution no resolved? yes problem closure major problem yes major problem review no end lessons learned review results service knowledge manage- ment system continual service improvement improvement actions communication
6 best practices: problem prioritization context with incident management problems are prioritized in low, medium, high and major/critical using the same structure and matrix as in incident management. details: for correct problem ticket prioritization the following rules shall apply: the incident priority is the input parameter from incident management. the event risk has to be evaluated within the problem management -for example in case of problems triggered by major incident make sure that the priority of the problem is major problem (priority 1). the risk of incident reoccurrence has to be evaluated within problem management. if the risk of incident reoccurrence is not known, use the risk level ‘normal’, if it’s possible to make an estimation, use ‘critical’ for problems with a high risk that an incident may appear for the same or other related cis. problem prioritization should be deduced from incident priority structure.
6 best practices: problem prioritization priority matrix priority of problem ticket proactive problems: in case the problem ticket is opened as a proactive problem, based on incident management or event data analysis, it is recommended to select problem priority medium or low unless there is a special reason to rate it higher. i n c i d e n t p r i o r i t y critical major problem major problem high major problem high medium high medium low medium high low normal risk of event reoccurance a major incident should always entail a major problem.
6 best practices: problem prioritization identifying the incident risk the following instructions are given as a guideline to find the right risk level. for this you need at least the following information: customer ci related services which can be disturbed if the given ci is crashed or damaged. (respectively to one or more customer) security information (if available) possible or coming work load or other information (e.g. external request) regarding to the ci or system. current maintenance information value risk of incident reoccurrence examples it is certain or very likely that the incident will occur respectively occurs again. the risk, that an incident may appear in the next few days, is high. it is possible and likely that the incident will occur respectively occurs again. the risk, that an incident may appear in the next few weeks, is predictable. a system is outdated and the maintenance needed doesn’t exist or is constrained. incident will appear due to an it can be foreseen, that an additional load or external effects (e.g. increased number of users or requests for a system, etc.). a security risk occurs and the potential risk that someone enters a system is high. an implemented workaround operates stable but isn’t a substitute for a durable solution the root cause and the solution are known, but the solution is not implemented or it takes a long time for implementation. high normal
7 best practices: problem investigation & diagnosis adapted problem flow service desk event management incident management proactive problem management supplier or contractor incident management implement workaround yes workaround needed? problem detection handover from major incident no raise known error record if required known error database change management rfc yes change needed? problem logging problem categorization problem prioritization supplier management cms problem investigation and diagnosis review changes if these caused the incident mi rca review & approval no problem resolution no resolved? yes problem closure major problem yes major problem review no end lessons learned review results service knowledge manage- ment system continual service improvement improvement actions communication
7 best practices: problem investigation & diagnosis presentation of root cause analysis to senior management for major problems major problems require it’s root cause analysis (rca) to be presented in a report to senior management of the service provider. our recommendation is to perform this in a weekly manner until there is no rca outstanding. after the final rca has been identified & accepted, a formal signoff is conducted. tracking: during rca investigation, the status is provided via a frequent report (2-5 times per week). it contains important and actual information about ongoing rcas including: root causes found root causes still under investigation incl. status and issues identified risks scheduled/planned de-briefings / sign off detailed streams and status of root cause analysis major problems require rca signoff from senior management
7 best practices: problem investigation & diagnosis overview workflow root cause analysis start rca check alarming chain check incident process key players during incident special topics other weaknesses incident after change no yes deep dive with chm identify root causes root cause classifications request rca from supplier responsibility supplier calculate effort for claim mgmt customer service provider reoccuring incident no fill known error claim mgmt determination of final incident priority & downtime define solutions to avoid recurrence get rca approval from involved parties rca template rca finished
7 best practices: problem investigation & diagnosis details workflow root cause analysis 1/5 check alarming chain check incident process when was mod service informed? explanation for delay of mod service activation. did a monitoring system detect the impact/did the customer detect the impact? was a standard checklist used to solve the incident? was the customer business impact clear during the whole incident process? has the responsible sdm been involved in clarification? evaluation of downtime, time to repair for this case; further explanation (optional) quality of incident details? was critical landscape used and sufficient? has configuration management data been sufficient? was a transition or transformation activity causing the incident? what went well / wrong? has a change caused the incident? (based on “check of changes”) key players during incident were all key players during incident management process available in time? which roles were missing? have required 3rd parties joined in incident resolution? which ones? did the 3rd party react within sla / ola?
7 best practices: problem investigation & diagnosis details workflow root cause analysis 2/5 incident cau- sed by change (deep dive with change) was incident really caused by a conducted change? was the change tested before implementation? change type (= change classification: major, significant, minor, standard) was change discussed in the appropriate change advisory board (cab)? date of change discussion at cab? was change approved by cab? was a backout method defined? did the defined backout method work as planned? if not, why not? was it a customer driven change? planned change start/end time? actual change start/end time? has the run book been followed or where have there been variances? was the run book reasonable and feasible ?
7 best practices: problem investigation & diagnosis details workflow root cause analysis 3/5 supplier involvement identify root cause root cause classification responsibility for incident which 3rd party suppliers were involved? do suppliers agree with the so far identified rca idea? was the root cause found? detailed root cause description use some of the techniques to identify the root cause as outlined in section 4.4.4.3 (itil service operation book). classify the rca like (hardware, software, application, etc…) identify responsibilty for the occured incident based on facts (no fingerpointing)
7 best practices: problem investigation & diagnosis details workflow root cause analysis 4/5 reoccurring incident fill known error database conclude final business impact and final downtime define solu- tions to avoid reoccurrence was this a reoccurring incident? is this incident relevant for other systems of the same customer? is this incident relevant for other customers/environments? check known error database for entry create/update entry in known error data base determine final business impact determine final start/end of time of impact and final downtime description of measures should include deliverable(s), responsible and due dates
7 best practices: problem investigation & diagnosis details workflow root cause analysis 5/5 get rca approval get agreement from responsible resolution measure owners gain acceptance for final rca from involved service delivery manager production responsible customer involved suppliers sign off rca at least for major incidents and important high incidents, introduce the rca to senior management in a sign off call to get the final approval
7 best practices: problem resolution adapted problem flow service desk event management incident management proactive problem management supplier or contractor incident management implement workaround yes workaround needed? problem detection handover from major incident no raise known error record if required known error database change management rfc yes change needed? problem logging problem categorization problem prioritization supplier management cms problem investigation and diagnosis review changes if these caused the incident mi rca review & approval no problem resolution no resolved? yes problem closure major problem yes major problem review no end lessons learned review results service knowledge manage- ment system continual service improvement improvement actions communication
8 best practices: problem resolution solution / measures tracking after problem management has an agreed rca, the identified measures are taken over into solution/measure tracking. a weekly tracking has to be organized by problem management to ensure that measures are performed as expected in scope and time. the problem manager ensures that each measure owner reports the actual status and informs about potentially overdue measures. responsibility: responsibility can not be outsourced! the service provider owns the decision to implement resolution measures. in case it has formally been decided to not perform recommended resolution measures, this decision should be documented in the corresponding known errors
9 best practices: problem closure adapted problem flow service desk event management incident management proactive problem management supplier or contractor incident management implement workaround yes workaround needed? problem detection handover from major incident no raise known error record if required known error database change management rfc yes change needed? problem logging problem categorization problem prioritization supplier management cms problem investigation and diagnosis review changes if these caused the incident mi rca review & approval no problem resolution no resolved? yes problem closure major problem yes major problem review no end lessons learned review results service knowledge manage- ment system continual service improvement improvement actions communication
9 best practices: problem closure solution / measures tracking after the implementation of all measures has been reported as complete, the problem manager performs a final quality check and closes the problem ticket in the ticketing system.
abbreviations & terms abbreviation explanation cab ci itil lim mi mod ola rca sdm sla change advisory board configuration item information technology infrastructure librarytm lead incident manager major incident manager on duty operational level agreement root cause analysis service delivery manager service level agreement