process training industry best practice incident management
agenda 1 2 3 4 5 6 7 our goal purpose of incident management general recommendations general role description incident flow best practices: incident prioritization best practices: major & high incident handling 8 9 10 11 best practices: investigation & diagnosis best practices: link to change management best practices: incident closure best practices: safeguarding
1 our goal itil is the de facto standard for the whole it industry. however the itil release cycles vary from 5-10 years and do not reflect the speed of business changed today. our association focuses on cross-company collaboration to define best practices based on our design and operational experiences. our goal is to provide practical industry best practices beyond the itil standards. our chapter for today : incident management!
2 purpose of incident management itil quote: “the purpose of incident management is to restore normal service operation as quickly as possible and minimize the adverse impact on business operations, thus ensuring that agreed levels of service quality are maintained. ‘normal service operation’ is defined as an operational state where services and cis are performing within their agreed service and operational levels.”
3 general recommendations information is everything make it easy to access required data for all participants of the incident process. document and share high level architectures of infrastructures and services. have a configuration database available as delivery components need to be documented. this database should also show which suppliers provide certain services for components or services. have a definition in your configuration database about the criticality of a service or device. this will significantly speed up detection of a high or critical incidents and therefore also speed up the alarming chain and thus the resolution. have predefined data collection mechanisms defined with your suppliers which are automatically triggered internally to collect logs or other data the suppliers need to provide fast and targeted support. have documented and tested for all different involved suppliers which data, from which location, via which method shall be collected and submitted to enable a faster resolution. have a solution defined and tested to enable retrieval of potentially large logfiles (gb size). the speed of the incident resolution process is mainly determined by the availability of information. so it’s of utter importance to have a good concept for information collection and documentation in place in which your supplier landscape is actively integrated.
4 general role description service delivery manager (sdm) operations manager problem manager problem analyst main interface to customer allocates all customer requests into the internal organization. responsible for contract fulfill- ment regarding delivery and commercial aspects plans, controls and is respon- sible for the adequate provision of one or a number of services has operative responsibility for the whole lifecycle of the problems main contact for the sdm provides the technical inventory and service data for the part of the service ensures required approvals (root cause analysis results, resolution actions, etc..) leads the root cause analysis including detailed documen- tation and involvement of external parties ensures of all necessary resources and of all required information sources change manager change coordinator change approver change requester analyzes and reviews rfc and change planning coordinates changes and is responsible for change planning is responsible for safeguarding of complex and high risk changes approves changes; leads cab ensures mandatory approvals (e.g. presents the changes to man- datory change advisory boards ) is responsible that change implementation is compliant with approved change planning e.g. sdm, cabs, customer, etc. reviews the change planning approves or denies the change e.g. customer, product manager, project manager, sdm, etc. requests a change with a de- scription of the cause and the target as a rfc (request for change) change implementer incident manager lead incident manager (lim) manager on duty (mod) is responsible for activities as part of a change implementation ensures processing of incidents during their entire life cycle within service level or operational level agreements § overall responsibility for the execution of the inm process responsible to handle escalations steers as a project manager the incident solution is responsible for the document- ation and handover to problem management supports incident solution with dedicated customer or service line knowledge ensures that incident solution will be pushed forward 24x7
5 incident management flow based on itil event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
6 best practices: incident prioritization event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
6 best practices: incident prioritization basics three basic acknowledgments: 1. the service provider decides the priority of an incident for all activities around the incident. 2. the service provider decides the pace via the incident priority. 3. suppliers need to understand the provider priority if it diverts from their standard. priority codes: for prioritization of incidents usually a matrix is used. while itil uses a 3 x 3 matrix resulting in 5 priority codes, we suggest to use a minimum of four priority codes. some organizations have more than 10 priorities defined. usually additional priorities beyond 5 just make things more difficult and do not result in noteworthy handling improvements. we recommend the following naming similar to itil: priority code 1 2 3 4 classification major incident (critical) high medium low procedure top management engagement (top) management engagement standard incident procedure standard incident procedure
6 best practices: incident prioritization priority classification critical critical (mi) high high medium medium medium medium high medium low high high high high medium medium medium medium medium medium medium medium critical medium medium medium medium medium medium low high medium low low low impact mi a critical custome r service is completel y down watch out ! known issues about classifications: “system is slow” this can cause a real headache for many organizations with their service suppliers. often sla do not differentiate between service is down and service is degraded. as a result suppliers are keen to determine if the system generally is still working or really down. we recommend to have a clear definition with your suppliers for slow cases also. standard incident procedure (low & medium): a non-critical service is down or has performance problems. low and medium ranked incidents follow the standard incident procedure: they are resolved within the responsible service unit via the ticketing system. high incident procedure: a critical service chain is partly out of service, has performance problems or has lost it’s redundancy. these incidents are resolved according to the high incident procedure which requires additional communication to high management and eventually to top management. the procedure is similar, but not the same as the major incident procedure. critical or major incident (mi) procedure: the highest priority of resolution is a critical or major incident. a critical service chain is completely out of service with no possible workaround. usually the whole company is impacted or the most important service(s). many suppliers align with this minimal specification. therefore in case incident tickets need to be exchanged, the mapping of priorities is straightforward. if the categories do not match up, a mapping is required by both partners causing effort, delay and cost.
6 best practices: incident prioritization how to decide the priority to decide that an incident is potentially high or critical, service providers (internally or externally) should have an agreed specification with their customers about the criticality of services. it is additionally recommended to also indicate how many employees or customers are impacted if a certain system goes down. a structured documentation which services (systems and applications) of the customer are critical ones is also a very important point (tool: critical landscape or at least as minimal documentation a spreadsheet) .
6 best practices: incident prioritization tools: critical landscape a critical landscape should be mandatory for all key customers. it is a summary of all critical components which are part of the delivery of a critical service per customer which is operated by the services provider. it should contain information like the following: most important components timeframe for criticality: what is the daily, monthly or annual timeframe when the services becomes critical, for example financial systems can have different critical timeframes like year end closing (critical) or mid month (not critical). contact persons for each system critical landscapes must be updated regularly, at least quarterly and has to be completed by the responsible service delivery manager or service operations manager. this document provides the basis to perform incident verification and prioritization. best practices show that the business impact should be communicated to involved persons in practical examples to illustrate the impact on the customer’s business. for example an out-of-service printer at the goods shipping center could cause a full stop of all goods delivery, or without goods withdrawal from the warehouse could stop the whole production line. a critical landscape is a description of the service chain with all critical components which are part of the delivery of a critical service for a customer. it offers a clear overview of the business impact for a customer in case a critical service is down.
7 best practices: major & high incident handling event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
7 best practices: major & high incident handling simplified, general incident workflow new incident critical / special case verify business impact medium or lower incident priority & quality check by service delivery manager or lead incident manager with account responsible / customer responsible mi procedure high procedure standard inm process mi procedure started critical or special case high lower than high • after first verification & analysis: -confirmed: start mi procedure -high procedure -lower than high: follow standard incident process • mandatory participants to ensure quick and qualified analysis: - lead incident manager - service delivery manager, service owner, production - technicians of affected service responsible regular status updates in mgmt call and via info mail + incident report regular status updates via incident report updates within ticket tool status mail will be sent after each management call
major incident 7 best practices: major & high incident handling major incident procedure major incidents are rare and have dramatic impact. therefore these are usually also of complex nature and can often not be resolved by simply rebooting a server. important is, that major and high incidents are detected quickly and handed over to a specified expert organization for evaluation and resolution. therefore local organizations shall spend maximum 45 minutes to identify, verify and decide if an incident needs to be handed over to the centralized organization via the “red phone”. the verification should be done by a customer service delivery manager or local lead incident manager. our recommended 45 minutes timeboxing is to prevent from situations that local teams try to solve an incident for an extended period of time without involving additional required experts or suppliers. with that the overall resolution time will likely be expanded. plenty may know from experience sentences like “only 10 more minutes and we will have it solved” or “just one reboot and the problem is solved”, but often this is not the case and the incident lingers on. after handover a mandatory, communicated, trained and tested process, with all involved parties should be implemented. characteristics of the major incident procedure: managed by the “red phone” as an central authority with skilled employees, 24/7 availability and the required expertise to take care of incidents of the highest priority partners and suppliers are involved in regular status calls to make use of their expertise mandatory check of all changes during the last 7 days are being performed full layer check: all cis and components of the affected service are checked using checklists and instructions to ensure that all technical issues are being detected involves senior management as well as the manager on duty continuous customer communication regular update calls on all counter steps and results
7 best practices: major & high incident handling red phone alarming workflow for major incidents – general major incident local organization local organization central organizationcentral organization central organization red phone scope new incident incoming call working for incident solution open / update / close final report / follow up activities full layer check completed full change check completed steps verification of the incident as a critical incident priority critical combined tech. & mgmt call is in place combined tech. & mgmt call mgmt status call mgmt status call safeguarding conf call / prm warm handover final incident report distributed 30 min after last call continuation as technician call until incident is solved supplier involvement supplier ticket has to be opened by line organization / operating teams escalation to key suppliers done by manager on duty escalation to other suppliersbased on contact matrix done by line organization mi review within 48 hours timetable participants customer information 45 min 20 min 20+x min 30 min 60 min manager on duty lead incident manager, service & operation responsible, technicians, supplier customer communication by service delivery management (informs and / or involves also account management / sales) must use incident report / checklist to analyse and solve incidents, and to communicate in the defined and mandatory structure / info mail documents incident report / weekly / daily status, change list, critical landscapes, architecture pictures, previous rca´s, known errors t n e m e g a n a m m e l b o r p o t r e v o d n a h
7 best practices: major & high incident handling activities and responsibilities within mi incident management process major incident yellow phone red phone lim (lead incident manager) technician lead mod (manager on duty) the yellow phone is also part of a centralized organisation. it is the single point of contact for all high incident´s and has the following responsibilities: operating as spoc providing complete communication infrastructure (conference calls, desktop sharing) internal documentation about involved teams or contacts incl. their handovers attends and making notes on tech. bridge the red phone consists of a centralised organisation, providing the overall lead incident management. it is the single point of contact for all major incidents and has the following responsibilities: operating as spoc providing complete communication infrastructure (conference calls, desktop sharing) internal documentation about involved teams or contacts incl. their handovers attends and making notes on tech. and mgmt. bridge performing supplier involvement and if required escalations preparing major incident notifications the lead incident manager is responsible for managing conference calls and ensures that the procedure is followed. will be named in the combined tech. & mgmt call during the start-up phase (usually a dedicated, customer lim or technical mod) coordinates all activities of the involved technical or operational teams (internal + external) and involved suppliers pushes the solution process from technical point of view documentation and reporting of the findings and implemented measures (incl. their results or effects) moderates and structures combined technician and management call steers and moderates the management status call responsible for continuous incident documentation (info email, incident report or action item list) ensures enablement for an ongoing technical analyses. takes care that the incident is handled according to the major incident guidlines setup of proper safeguarding ensures smooth handover to problem management– preparing root cause analysis template. covers dedicated work streams supports incident solution with dedicated customer or service line knowledge operates or steers dedicated technician conference call on request ensure supplier escalation responsible for people management makes requiered decisions sdm (service delivery manager) evaluates the priority and establishes a continuously customer communication sdm represents the interface to the customer covers dedicated work streams
7 best practices: major & high incident handling yellow phone alarming workflow for high incidents high incident local organization local organization central organizationcentral organization central organization scope new incident incoming call working for incident solution open / update / close final report / follow up activities yellow phone steps verify incident priority as high initiate high incident procedure and involvement mod full layer check completed full change check completed 60 min continuation as technician call until incident is solved supplier ticket has to be opened by line organization / operating teams manager on duty lead incident manager, service & operation responsible, technicians, supplier customer communication by service delivery management (informs and / or involves also account management / sales) supplier involvement participants customer information must use incident report / checklist to analyse and solve incidents, and to communicate in the defined and mandatory structure / info mail documents incident report / weekly / daily status, change list, critical landscapes, architecture pictures, previous rca´s, known errors t n e m e g a n a m m e l b o r p o t r e v o d n a h
7 best practices: major & high incident handling activities and responsibilities within hi incident management process high incident yellow phone red phone lim (lead incident manager) technician lead mod (manager on duty) the yellow phone is also part of a centralized organisation. it is the single point of contact for all high incident´s and has the following responsibilities: operating as spoc providing complete communication infrastructure (conference calls, desktop sharing) internal documentation about involved teams or contacts incl. their handovers attends and making notes on tech. bridge the red phone consists of a centralised organisation, providing the overall lead incident management. it is the single point of contact for all major incidents and has the following responsibilities: operating as spoc providing complete communication infrastructure (conference calls, desktop sharing) internal documentation about involved teams or contacts incl. their handovers attends and making notes on tech. and mgmt. bridge performing supplier involvement and if required escalations preparing major incident notifications the lead incident manager is responsible for managing conference calls and ensures that the procedure is followed. will be named in the combined tech. & mgmt call during the start-up phase (usually a dedicated, customer lim or technical mod) coordinates all activities of the involved technical or operational teams (internal + external) and involved suppliers pushes the solution process from technical point of view documentation and reporting of the findings and implemented measures (incl. their results or effects) moderates and structures combined technician and management call steers and moderates the management status call responsible for continuous incident documentation (info email, incident report or action item list) ensures enablement for an ongoing technical analyses. takes care that the incident is handled according to the major incident guidlines setup of proper safeguarding ensures smooth handover to problem management– preparing root cause analysis template. covers dedicated work streams supports incident solution with dedicated customer or service line knowledge operates or steers dedicated technician conference call on request ensure supplier escalation responsible for people management makes requiered decisions sdm (service delivery manager) evaluates the priority and establishes a continuously customer communication sdm represents the interface to the customer covers dedicated work streams
7 best practices: major & high incident handling tools: combined technician / management conference call combined technician/management conf call (hosted mod): major incident the combined call of management and technicians initiated by the red phone is the initial and only platform to coordinate and assign all technical measures into work streams, including sharing the results. the call will be initial steered by lim, to make the initial evaluation of the situation and set-up further activities. furthermore he assigns a responsible tech lead/s (in general this will be the responsible mod, a customer specific mod for customer or a dedicated lim). this depends on the circumstances and who is impacted. basically the best available and trained person for the job needs to be assigned. the call starts with all necessary participants – management representatives will be released after initial information and structure setup into management status calls. in case the business line, account or customer has already established a technician call, the further procedure will be aligned in the combined technician / management call. immediately after the management team left the combined call, the technicians continue in the same or a separate permanent open call to work as a team first on predefined work stream to perform a “layer check” and a “check of changes”. further work streams will be defined, depending on the first findings and the affected technology for the technical drill down. during the whole incident, the assigned lim ensure the communication with structured list and status email. all required documents, call setup or collaboration platform, etc. will be provided by the mod or lim. general participants: lead incident manager manager on duty technical experts l2/l3 sdm (customer related) service owner & operation responsible supplier key roles (all needed mods) top management team the combined technician / management conference call is the platform to coordinate and to assign all technical measures. after initial management information a separation of management representatives and technicians is recommended to get a faster incident resolution.
7 best practices: major & high incident handling tools: info mail – red phone incidents major incident the info mail is a fast and shortly summarized incident description which receives an update after each call. the info mail is mandatory for every case steered by a lim. it is not replacing the official documentation (incident report). it is always based on the previous status mail and has to be written and send out as soon as the lim has enough details about the ongoing incident, but latest for the first time 15 minutes after the combined technician and management call has started. the email is sent out to a pre-defined distribution list by the lim directly. first mail: contains information such as: customer, priority, impact on customer, initiated actions, next updates update mail: contains information as: changed impact on customer, changed priority, update/ results of actions including supplier escalations, next steps, next update time mail subject: reflects the priority situation from customer perspective and stays always the same until priority got lowered or the incident is solved. if the incident is closed and the service is still in the safeguarding mode, a final mail with subject line: “final: customer -service restored –needs to be sent out. the info mail is an important instrument to provide clear status information to all relevant stakeholder.
supplier involvement
7 best practices: major & high incident handling supplier involvement overview mi procedure high procedure standard inm process generating supplier support cases (‘tickets‘) in general has to be done by the responsible delivery team according to the standard incident management process generating supplier support cases (‘tickets‘) in general has to be done by the responsible delivery team according to the standard incident management process key supplier mod has to initiate key supplier escalation according to the agreed governance model line organization (respective operating teams) has to escalate based on the contact matrix of the line specific contact map (yellow phone has to escalate based on the contact matrix of the line specific contact map)
8 best practices: investigation & diagnosis event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
8 best practices: investigation & diagnosis basic questions & full layer check basic questions: before starting any steps to resolve the incident, a brief overview about the situation should be establis- hed. according to the itil framework the following 5w questions should be answered: what happened? who did that? when did it take place? where did it take place? why did that happen? method: full layer check always check all layers unless a layer does not apply -> example right side this check is a fast rough check to identify in which layers issues exist after the full layer check we dive into layers with issues full layer check status of the layer check: responsible person notes: details relevancy check of data center infrastructure: check of network & its components (incl. firewalls): check of application: check of server and operation system: check of backup & restore: check of storage (san, nas, filter): check of database and middleware: check of application and job processing: a structured diagnosis expedites the incident resolution.
9 best practices: link to change management event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
9 best practices: link to change management check of changes a check of the recent changes during the last 7 days prior a high or major incident is important as often the incident occurs at this point in time due to a change in the infrastructure. the mod will secure the execution of the change check and will assure that all relevant information will be available for the lead technician for inclusion in the technician call. the technical lead has to ensure a relevance check of all changes against the current incident situation. results of change verification will be shared in the second management call the start and end time, as well as status for the “check of changes” activity has to be documented in the trigger of events in the incident report. a structured check of changes is mandatory in case of major or high incidents.
9 best practices: link to change management emergency changes during incident diagnosis or afterwards changes to the infrastructure might be required. we are strongly suggesting not to perform quick updates, patches, reboots or similar changes to the infrastructure as these often decrease the chance to indentify the real root cause of the issue. as a result our recommendation is to use an emergency change procedure in which all changes to the infrastructure are reviewed by the mod and authorized prior application. for this authorization the change number must exist in the regular change management system as well as at least a short e-mail based documentation of what is intended to be changed. change documentation in the change system has to be performed after the incident is being resolved within a specified timeframe (recommendation is maximum 1 working day) the emergency change procedure offers a structured way to implement changes during ongoing incident resolution activities.
9 best practices: link to change management mod change handeling: emergency changes & short leadtime changes emergency changes definition not in scope approval documentation emergency change short lead time change a short lead time change cannot fulfill the regular the process due to several valid reasons. lead time as defined in an emergency change has to be carried out immediately in order to resolve an incident (without an incident there cannot be an emergency change). emergency changes are meant to avoid massive damages and have to be authorized at least by an mod (manager on duty) and the emergency change advisory board. late change planning does not justify an emergency change. changes without a relation to a incident situation emergency changes have to be approved by mod and in case customers are affected by the impacted customer. short lead time changes have to be approved by mod and in case customers are affected by the impacted customer. the detailed documentation in an emergency change ticket may be done retrospectively when time is of the essence (latest at the following business day). however, a change number and short abstract of the change must exist for gaining approval. the emergency change must be related with the incident to be solved. a short lead time change is always completely documented before change implementation. the change requires a documented justification for the short lead time.
10 best practices: incident closure event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
10 best practices: incident closure final incident priority the priority of incidents can change from critical to low over one business day. however the highest verified priority within the lifetime of an incident has to be documented before closure. different behavior means falsifying/manipulating the incident statistics and potentially inhibiting the handover to problem management. incident is found medium since no one is currently using the service and only partly unavailable incident becomes critical because service breaks down completely incident gets solved medium critical critical on closure, in this case a priority of „critical“ has to be documented! morning noon evening high low impact raised since more end users start using the service effectively impact significantly reduced since service mostly restored and end users went home
10 best practices: incident closure major incident (mi) review a major incident review is necessary after every major incident or on special management request. the review normally takes place after the last call, together with all technical and management key players of the related incident and is hosted by the mod and moderated, structured by a lim. basis for the mi review is the „final incident report“. target is to review the complete incident history and is focused on: - are the trigger of events correct and complete? - what led to the solution of the incident? - information for a preliminary root cause or additional information relevant for problem management - focus areas for problem management - identified weak points during the major incident process - people/teams required for problem management all topics will be recorded in a document/ticket for problem management. the major incident review has to be done during office hours from monday to friday and the assigned problem manager has to attend the process. the major incident review is important to get an overview what went good or bad during the incident management process and to ensure that all required information will be provided to the problem management.
10 best practices: safeguarding event management web interface phone call email incident identification major incident procedure yes major incident procedure no initial diagnosis key best practice content available is this really an incident? no to request fulfilment (if this is a service request) or service portfolio management (if this is a change proposal) functional escalation? yes functional escalation? yes escalation needed? no no yes incident logging incident categorization management escalation? yes hierarchic escalation? no change mgmt major incident procedure yes major incident procedure adapted from itil framework: figure 4.3 incident management process flow adapted by zero outage industry standard no investigation and diagnosis resolution identified? yes resolution and recovery incident closure safeguarding end incidentpriorization
11 best practices: safeguarding basics after resolving a major or high incident, services are restored. due to the high impact we recommend to perform a safeguarding phase . in this phase disturbed services are being specifically monitored with high attention for a specified period. safeguarding measures and follow-up activities are defined by the lim in collaboration with the participants of the management conference call. the safeguarding method (extended or standard/default) depends on a risk assessment. due to the high impact of major and high incidents it is required to set up safeguarding measures after the resolution of the incident.
11 best practices: safeguarding risk assessment matrix to define safeguarding measures risk error detection and correction safeguarding method very high incident disappeared without conscious executed action. incident cause is completely unknown. incident mode remains open until the risk of reoccurrence can be rated as max. high. high incident has been solved. but incident cause has not been identified for sure (e.g. after simple reboot of a component). default safeguarding + additional safeguarding e-mail which should contain a detailed description of the taken safeguarding measures, incl. planned actions in case of reoccurrence, a defined distribution list, telephone conference templates and special on-call duty overview. medium incident cause has been identified and stable workaround fixed the problem. safeguarding measures if deemed necessary. low incident cause has been identified without any doubt and is already fixed. no safeguarding required.
abbreviations & terms abbreviation explanation cab ci inm itil lim mi mod prm sdm sla spoc change advisory board configuration item incident management information technology infrastructure librarytm lead incident manager major incident manager on duty problem management service delivery manager service level agreement single point of contact