You are using an unsupported browser. Please update your browser to the latest version on or before July 31, 2020.
You are viewing the article in preview mode. It is not live at the moment.
Home > Company Policies > Information Technology > Data Systems Downtime Policy
Data Systems Downtime Policy
print icon

Purpose

  • To establish a backup process in the event of data system failure at our DCs or Medical Fitness Centers.
  • To outline the general procedures that end- users should follow to ensure continuance of patient care during data systems downtime.
  • To define the enterprise-wide downtime communication process for all the clinical and administrative staff in a timely, discrete, and secured manner.

Policy Statement

  • Although Data systems should be operational all the time Data Systems might have some unexpected failure, which might result in disruption in the provision of patient care services.
  • This Policy describes what actions should be taken at the time that an unexpected failure happens, to continue offer quality service.
  • This policy addresses how to react against the situation of software- or hardware-related malfunctions only, not against natural disasters.
  • This policy is applicable to all the Clinical and Administrative Staff of Smart Salem Medical Centers who are working with Data Systems.
  • This policy is to provide an end user support guideline for downtime situation at SSMSC.
  • SSMSC is committed to ensuring reliable information technology services. The goal of this policy is to explain those circumstances during which downtime may occur, anticipated durations of downtime events, and procedures for notifying affected users.
  • SSMSC is committed to ensuring reliable information technology services. To meet this objective, SSMSC systems may need to be taken offline to maintain or improve system performance, safeguard data, or to respond to emergency situations.

Responsibilities

Clinical Staff

The detailed responsibilities of clinical staff, department heads and the medical records team are described in the later session. Each clinical team member is responsible for following the procedures during and after downtime. This is a mandatory policy to be followed.

IT Team

  • To operate and maintain HIS/CORE SmartSalem System in healthy condition.
  • If technically possible, to implement configuration of clusters in order to eliminate any of singular point of failure (SPOF) for the important systems such as HIS/ CORE SmartSalem System to the phased implementation plan as defined in the IT Operations Plan
  • To perform regular maintenance activities regularly to minimize the possibility of malfunctions.
  • To prepare for a file sharing system which can share the clinical document templates at the time of a system failure on HIS/LIS
  • To prepare for the Backup Plan, describing how to perform regular backups in the way to minimize data loss in case of a system failure.
  • To prepare for the Recovery Plan, describing how to perform system recovery activities when the system has been determined as unrecoverable from the production servers and necessary to recover from the backup images and/or backup data.
  • To determine whether it is a system failure or not when end-users report malfunctions related to HIS/ CORE SmartSalem System and Data Systems
  • To report to CTO as soon as a system failure has been determined, along with the expected information on the recovery timeline.
  • To report to CTO as soon as HIS has been restored and confirmed as working normal.
  • Planned downtime to be recorded as change request. All the downtime information shall be maintained.
  • Unplanned downtime shall be registered as “incident”, with all supporting documents, in the service desk system.

Procedure

To Determine a System Failure

  • End-users of Data Systems should report IT related issues to IT Team (IT Director or onsite IT Support staff).
  • IT Applications and Infrastructure Managers should ensure that their Staff including Applications Specialist, Systems Administrator and Network Engineer are monitoring the IT Systems (Infrastructure and Applications) and to report to IT Director whenever they discover any kind of failure which could disrupt the normal operation of any data systems.
  • When IT Helpdesk Engineers or Application Specialists get calls from end-users regarding malfunctions of any data system, which cannot be fixed by normal procedures as they daily do within an apparently short time, they have to check with the IT Applications Manager/IT Infrastructure Manager regarding the possibility of a system failure on the Data system.
  • Based on overall information from end-users and IT staff, IT Director should determine whether the data systems have a system failure or not
  • Once the Data system has turned out to have a System Failure, IT Director should determine whether it is a Partial System Failure or Total System Failure
  • In case that all the modules of the Data System could not be operational at all departments, this shall be regarded as a Total System Failure
  • When some, but not all, modules could not be operational, this shall be regarded as a Partial System Failure
  • In case that the Data System could not be operational only at some limited areas or departments, this shall be regarded Partial System Failure
  • IT Director should communicate instantly with the IT Applications Manager and Infrastructure Manager to setup the Recovery Action Plan describing how to perform the system recovery activities, even if Data System has been determined as unrecoverable from the production servers and necessary to recover from the backup data and/or backup images.
  • The Recovery Action Plan should be ready instantly based on the standardized procedure defined in the IT Disaster Recovery Plan
  • Once the Recovery Action Plan is ready, IT Director has to report to CTO promptly with the relevant and up-to-date information including the expected recovery timeline.
  • Entry of data into the manual entry forms should be initiated by the respective department heads or allocated supervisor within the department.

It’s the responsibility of the department heads and individual staff to ensure sufficient manual forms are always available for all the users in the department, for all the service providers who are directly or indirectly involved in providing care in the unit (CHO, Medical Staff, Back office etc).

Emergency Downtime (Unplanned Downtime)

Unexpected circumstances may arise where systems or services will be interrupted without prior notice. Every effort will be made to avoid such circumstances. However, incidences may arise involving a compromise of system security, the potential for damage to equipment or data, or emergency repairs. If the affected system(s) cannot be brought back online immediately, affected users will be contacted via the Notification of Downtime mechanism described below.

Planned Downtime

From time to time, it will be necessary to make systems unavailable for the purpose of performing upgrades, maintenance, or housekeeping tasks. The goal of these tasks too is to ensure maximum system performance and prevent future system failures.

The following activities fall within the definition of Planned Downtime:

  • Application of patches to operating systems and other applications to fix vulnerabilities and bugs, add functionality, or improve performance.
  • Monitoring and checking of system logs.
  • Security monitoring and auditing.
  • Disk defragmentation, disk clean-up, and other general disk maintenance operations.
  • Required upgrades to system physical memory or storage capacity.
  • Installation or upgrade of applications or services.
  • System performance tuning.
  • Regular backup of system data for the purpose of disaster recovery.

In the event that any of these activities will require downtime to perform, every effort should be made to perform the procedure during off-hours in order to minimize the impact on those who use the affected systems or services.

The following time periods will be used to carry out Planned Downtime activities:

  • Before 6:00 AM
  • After 10:00 PM

On occasion, it may be necessary to have Planned Downtime during regular business hours, namely if outside personnel are required to perform more elaborate procedures. If this is the case, then this Planned Downtime will be communicated to identify users of affected resources using the Notification of Downtime mechanism described below.

Notification of Downtime to End Users

Relevant Users will be notified of downtime according to the following procedure:

  • The applications specialist for the concerned data system is responsible for preparing the appropriate notification message with all required details for Planned Downtime, as well as any unplanned interruptions (emergency downtime) to data system availability as they occur.
  • The application specialist will notify IT service desk officers immediately, providing the required details of the affected system(s), type of downtime (Planned or Unplanned/Emergency), reasons for downtime, excepted duration and contact details of IT personal during the downtime.
  • The SSMC IT Helpdesk will first notify all affected users via email. All users are responsible for checking their email for downtime and system status notifications. In the event that email is unavailable due to Emergency Downtime, the helpdesk team will contact department heads by telephone to inform them of the situation or use Public Addressing System with Code Gray.
  • For Planned Downtime the Application Manager/ Service Desk Manager must give two (2) business days’ notice prior to the anticipated system unavailability. This step must be taken regardless of whether the downtime is scheduled to take place during off hours or regular business hours. Any exceptions should be approved by the IT Director.
  • In the event of Emergency Downtime, Service Desk Manager will use his/her discretion in notifying end users of the situation. In emergency circumstances where time is of the essence, it may not be possible for the Service Desk Manager to engage in normal downtime notification activities.
  • Application Manager/ Service Desk Manager will announce Unplanned Downtime based on investigation carried out by application owner in SKMCA IT team. When emergency measures are completed, or if 30 minutes has elapsed with no resolution, then the Application Manager / Service Desk Manager will team will contact all users with information on system status and/or information on additional expected downtime.

All downtime announcements will provide the following information:

  • Systems and services that are affected, as well as suggested alternatives to them (if any). For instance use of downtime form and manual clinical sheets.
  • Manual processes to be followed or any alternative processes for the service which is not available/accessible.
  • Start and end times of the Planned Downtime period, or estimated time to recovery in the event of Emergency Downtime.
  • The reasons why the downtime is taking place (in all the cases except in emergency downtime if the reason is not identified yet)
  • Any ongoing problems that are anticipated as a result of the downtime event.
  • All Issues reported or encountered during the downtime should be tracked and resolved via Service Desk application.

Post Downtime Procedures

  • IT will announce, once the service is back in operation.
  • The Maintenance of downtime forms is a responsibility of respective departments.
  • Entry of the data from the manual forms to the system should be the responsibility of individual department heads or assigned personnel within the department.
  • The IT / Systems manager should maintain the downtime log with the information on the number of announcements made via different mediums and the content which was announced.
  • Medical Records team / Department heads perform random audits to ensure that downtime forms are filled back in the system.

Clinical Systems Downtime Procedures:

CORE SmartSalem/HIS Downtime procedure & resumption checklist:

During Downtime Recovery
Registration Registration
MF Staff to use Salem Portal and Well Registration staff to manually capture the client details. Once the system is online, the patient file needs to be created in HIS using all the information available from the manual files including the manual file number given. MF to add the Manual application details in the dashboard
Orders Orders
Blood Collection Nurse to check in for MF clients on Salem Portal. NA
Radiology Staff to check in for MF clients on Salem Portal. NA
Laboratory Laboratory
Lab Staff to all sample details in the DHA System.  
Billing Billing
Billing team need to maintain a list of prices in manual copy and manual bill need to be raised. Once the system is online, the services needs to be entered in HIS and billed accordingly
IT Department IT Department
Notify all affected departments about the downtime, system affected, estimate downtime, and expected time for recovery. 1. Once the systems are recovered, the IT Helpdesk will notify the end-users with the system resumption. 2. IT will complete documentation with the following documentation: When did the downtime occurred Reason for downtime Communication actions taken. Update the downtime log.

Abbreviations and Definitions

HIS : Hospital Information System

System Failure : An event where an IT system could not be operational, except for the scheduled down time for regular maintenance.

Partial System Failure : An event that some, but not all, functions of an IT system could not be operational, or an IT system could not be working only at some limited areas/departments.

Total System Failure : An event where all functions of an IT system could not be operational at all departments in SSMC.

Single point of failure (SPOF) : A single point of failure is a part of a system that, if it fails, will stop the entire system from working. SPOFs are undesirable in any system with a goal of high availability or reliability.

End-User : A person who uses or is intended to ultimately use a product. The end user stands in contrast to users who support or maintain the product, such as system administrators, database administrators, or technicians.

Verification and Validation (V&V) : Independent procedures that are used together for checking that a product, service, or system meets requirements and specifications and that it fulfills its intended purpose.

Verification : The evaluation of whether a product, service, or system complies with a regulation, requirement, specification, or imposed condition. It is often an internal process. Contrast with validation.

Validation : The assurance that a product, service, or system meets the needs of the customer and other identified stakeholders. It often involves acceptance and suitability with external customers. Contrast with verification.

Backup Plan : Describes the process of backing up, which refers to the copying and archiving of computer data so it may be used to restore the original after a data loss event including following content.

Backup Scheme : Unstructured, Full, Imaging, Incremental, Differential, Reverse Delta, and/or Continuous Data Protection

Storage Media : Hard Disk Drive, Solid State Storage, Optical Storage such as CD, DVD etc.

Data Repository Management : On-line, Near-line, Off-line, Off-site data protection, Backup site or Disaster Recovery Center Selection and Extraction of Data.

Manipulation of Data and Dataset Optimization : Compression, De-duplication, Duplication, Encryption, Refactoring, Staging

Recovery Plan : A documented process or set of procedures to recover and protect a business IT infrastructure in the event of a system failure by utilizing backup data as well as system backup images, if necessary

Recovery Action Plan : The detailed action plan for recovery activities, which shall be set up based on the Recovery Plan and the actual situation at the time of a system failure.

Production Server : The production server is also known as live, as it is the server that users directly interact with

Test Server : To aid the software testing cycle in the software delivery process, it is essential to provide a validated, stable and usable test server to execute the test scenarios or replicate bugs

Feedback
0 out of 0 found this helpful

scroll to top icon