An IT service-level agreement should create clarity before pressure begins. When email is unavailable, a branch loses connectivity or a senior user cannot access a business system, nobody should be debating who to contact, whether the issue is critical or when the next update will arrive.
Many SLAs fail because they promise impressive response times but do not define business impact, service hours, dependencies or customer responsibilities. A provider may acknowledge every ticket within fifteen minutes while the same incidents remain unresolved for days. A business may label every request urgent, making genuine emergencies harder to identify.
A practical SLA is not a marketing promise. It is an operating agreement between the customer and support team.
Begin with the service, not the timer
Before response targets are written, define what the SLA applies to. The document should identify:
- supported users and legal entities;
- offices, branches, warehouses and remote workers;
- service hours and working days;
- supported devices, cloud services and infrastructure;
- included applications;
- support channels;
- onsite coverage;
- after-hours or emergency arrangements;
- third-party dependencies;
- excluded project and change work.
A response commitment has little meaning when the underlying scope is unclear.
Separate response, restoration and resolution
These terms should not be used interchangeably.
- Response time: how quickly the provider acknowledges, assesses and assigns the issue.
- Update interval: how frequently the customer receives progress information.
- Restoration time: how quickly a usable service or workaround is restored.
- Resolution time: how long it takes to complete the permanent fix and closure.
A network outage may be restored through a backup link while the faulty circuit remains under investigation. The business is productive again, but the root cause is not yet resolved. A good SLA recognises both events.
Use impact and urgency to determine priority
Priority should be based on the effect on operations and the time sensitivity of the issue.
| Priority | Impact example | Expected handling |
|---|---|---|
| P1 – Critical | Major business service unavailable, many users affected, serious security incident or confirmed recovery failure. | Immediate coordination, frequent updates, executive escalation and continuous effort within the agreed coverage window. |
| P2 – High | Important department or function unavailable with no acceptable workaround. | Rapid response, named owner, scheduled updates and priority restoration. |
| P3 – Normal | Individual or limited impact with a workaround available. | Handled within normal support workflow and business hours. |
| P4 – Low | Information request, standard access request, minor inconvenience or planned task. | Scheduled according to queue, approval and change requirements. |
The matrix should include examples from the customer’s environment. For a retail business, POS or branch connectivity may have high impact. For a professional-services firm, email, document access or secure remote working may be more critical.
Do not let job titles define priority
Executive users may require a specialised communication and support process, but priority should still reflect business impact. A chief executive’s minor printer issue should not displace a company-wide outage. At the same time, the provider should understand when an executive meeting, board presentation or customer event creates genuine urgency.
The SLA can include a VIP communication path without changing the underlying priority model.
Set realistic targets by coverage model
A remote-first support contract, an IT AMC with scheduled visits and a fully managed service are different operating models. Targets should reflect the people, tools and hours included.
When setting targets, consider:
- whether monitoring is active;
- whether support is remote, onsite or hybrid;
- engineer availability and location;
- after-hours coverage;
- customer approvals;
- hardware replacement and spare availability;
- internet, cloud and application vendors;
- building access and security procedures.
Promising a fixed resolution time for every issue is rarely credible. The SLA should define what is controlled by the provider and how external dependencies are managed.
Define the clock carefully
The agreement should state when SLA measurement starts, pauses and ends.
Questions to settle include:
- Does the clock begin when the ticket enters the approved system?
- Are messages to individual engineers counted?
- Does time outside service hours count?
- When does the clock pause for customer information or approval?
- How are vendor and hardware delays treated?
- When is a ticket considered restored or resolved?
- Who may reopen a ticket?
Without these rules, monthly SLA reports become arguments about data.
Use one controlled support channel
Users need an easy way to ask for help, but the organisation also needs a record. Approved channels may include a support portal, email address, telephone line or monitored chat channel.
Every request should become a ticket with:
- requester;
- affected service and location;
- business impact;
- priority;
- owner;
- timestamps;
- actions and communication;
- resolution and closure notes.
Informal messaging can remain useful for urgent communication, but it should not replace the ticket record.
Build an escalation ladder
Escalation should occur because impact is high, progress has stalled or a decision is needed—not because somebody knows a senior manager’s phone number.
A simple escalation ladder may include:
- service-desk owner;
- technical or functional specialist;
- service manager;
- customer IT or operations owner;
- executive sponsor for prolonged or severe disruption.
The SLA should list triggers such as missed update, repeated failed action, unresolved vendor dependency or risk of breaching a restoration target.
Define communication during major incidents
Technical work and customer communication should not compete. For a critical incident, assign an incident lead and a communication owner.
Updates should state:
- what is affected;
- when the incident began;
- business impact;
- current action;
- available workaround;
- dependencies or decisions required;
- time of the next update.
NIST’s current incident-response guidance places preparation, detection, response and recovery within wider cybersecurity risk management. The NIST SP 800-61 Revision 3 is a useful reference for building major-incident responsibilities that extend beyond the helpdesk.
Include security incidents without oversimplifying them
A suspected phishing compromise, ransomware alert or unauthorised administrator action may require urgent handling even if users can still work. The SLA should state how security incidents are identified, escalated and transferred to specialist response resources.
It should distinguish:
- routine security support;
- confirmed or suspected incident response;
- forensic investigation;
- customer, insurer, legal or regulatory communication;
- third-party security-service involvement.
A normal IT support contract may not include every specialist action, but the escalation path must be known.
Set service-request targets separately
A new user, application access, equipment request or mailbox change is not an incident. These requests often depend on management approval, HR information, licenses or hardware.
Define standard fulfilment targets for common requests, for example:
- new-user account setup after complete approval;
- access-group changes;
- license assignment;
- standard software installation;
- device preparation;
- employee offboarding;
- shared mailbox or distribution-list change.
The required information and approval should be part of the request form.
Control planned changes
Firewall changes, server patching, network reconfiguration and major Microsoft 365 policies can solve problems but also create risk. The SLA should identify which changes require:
- business approval;
- maintenance window;
- testing;
- backout plan;
- user communication;
- post-change verification.
Emergency changes should be documented and reviewed after the incident.
Account for third-party vendors
Internet providers, cloud platforms, application vendors and hardware suppliers can affect restoration. The support provider should still own coordination when it is within scope.
Measure:
- time to identify the dependency;
- time to open the vendor case;
- frequency of follow-up;
- customer decisions required;
- workaround effort;
- final root-cause record.
The provider should not be penalised for time it cannot control, but it should remain accountable for active coordination and communication.
Measure quality, not only speed
A provider can meet response targets while delivering poor service. A balanced scorecard should include:
- response and restoration performance;
- repeat incidents;
- tickets reopened;
- ageing backlog;
- customer satisfaction;
- major-incident reviews;
- preventive actions completed;
- backup and security exceptions;
- documentation updates;
- service-improvement commitments.
Fast closure is not useful when the same issue returns every week.
Use exclusions and customer responsibilities
The customer also affects service performance. The SLA should require:
- accurate user and asset information;
- timely approvals;
- valid licenses and warranties;
- access to premises and systems;
- use of approved support channels;
- notification of employee and infrastructure changes;
- cooperation during testing and recovery.
Provider exclusions and customer responsibilities should be written in plain language.
Review the SLA rather than freezing it
The priority model and targets should be reviewed when the business adds branches, critical systems or extended working hours. Monthly reporting may reveal that certain requests need a standard catalogue, that a recurring incident needs a project or that after-hours demand is higher than expected.
The service review should produce actions rather than merely present percentages.
The SLA design worksheet
- List supported services and business hours.
- Define P1–P4 using the customer’s real operations.
- Separate response, update, restoration and resolution.
- Set realistic targets for the chosen delivery model.
- Define when clocks start, pause and close.
- Choose controlled support channels.
- Create technical and management escalation paths.
- Define major-incident communication.
- Separate incidents, requests, changes and projects.
- Document vendor dependencies.
- Measure quality and recurrence, not only speed.
- Review service data and improve the SLA quarterly or when the environment changes.
Frequently asked questions
What is a reasonable IT support response time?
It depends on business impact, service hours and delivery model. Critical outages require rapid engagement, while routine requests can follow normal scheduling.
Should resolution times be guaranteed?
Guaranteed restoration may be possible for defined services and designs, but permanent resolution often depends on hardware, vendors, approvals or complex diagnosis.
Does the SLA apply to WhatsApp messages?
Only when the agreement defines it as an approved channel and ensures each message becomes a controlled ticket.
What should happen after a major incident?
The provider should document impact, timeline, cause, recovery actions and preventive improvements, with owners and deadlines.
How often should the SLA be reviewed?
Review it at least during regular service reviews and whenever locations, working hours, systems or business-critical processes materially change.
A useful SLA makes support predictable without pretending every technical issue can be solved to a fixed stopwatch. It defines ownership, communication and business priorities. For a wider support model connecting helpdesk, monitoring, Microsoft 365, infrastructure, security and reporting, review managed IT services for Dubai businesses.