Anyone who must select IT service providers decides not only on daily rates and availability, but also on an expansion of their own attack surface, on regulatory obligations to provide evidence, and on how quickly you can be operational again in the event of an incident. In many companies, service providers effectively become co-operators: they administer systems, move data, configure security mechanisms and influence operational processes. Precisely for this reason a ‚good impression in the pitch‘ is not sufficient.
This article is intended as a decision aid for IT management, information security (ISMS), data protection, compliance and executive management. It presents a practical approach to evaluate providers on a risk basis: which evidence is sensible, which controls are indispensable, how to establish governance and responsibilities, and how to avoid common pitfalls (subcontractors, log access, exit, chains of evidence). The goal is a selection decision that will withstand later audit and will not lead to surprises in operation.
1) Starting point: What does „risk“ mean concretely for IT service providers?
In the supplier context risk is not abstract but very concrete: it is the combination of likelihood and impact on confidentiality, integrity and availability (the CIA triad). In addition, verifiability counts: can you demonstrate that controls exist and are effective?
For the assessment it helps to clearly separate four risk domains:
- Data risk: Which data does the provider process (personal, confidential, trade secrets)? Where are they located, how are they encrypted, who has access?
- Access and operational risk: Does the provider receive admin access, shell access, VPN access, or only ticket-based access? Are there emergency access paths? Are changes performed under change management?
- Supply chain risk: Subcontractors (e.g. cloud providers, support chains), location changes, dependencies, proprietary tools, exit hurdles.
- Compliance and audit risk: GDPR, industry rules, internal policies, retention, logging, evidentiary obligations, inspection rights.
A sound selection decision requires that you first precisely describe the planned deployment of the provider: systems, data classes, privileges, operating windows, RTO/RPO (Recovery Time Objective/Recovery Point Objective) and organizational interfaces. Without this delimitation every assessment will either be too soft („everything is important“) or too strict („everything forbidden“).
2) Risk classification before screening: Tiering instead of gut feeling
In practice, a tiering (classification into criticality levels) that governs the effort for due diligence and contractual clauses proves effective. This prevents you from having to roll out the entire audit machinery for every minor support contract — and at the same time ensures that critical partners do not slip through.
Proposal for 4 tiers (adjustable)
- Tier 1 – Critical: admin/root access, networks close to production, processing of sensitive data, operation of critical processes, significant outage impact.
- Tier 2 – High: access to important systems or large data volumes, but with constraints (e.g. only via jump host, strictly separated scope).
- Tier 3 – Medium: limited access, predominantly advisory/project work, RESTricted data processing.
- Tier 4 – Low: no or minimal data processing, no access to productive systems (e.g. training).
Tiering should not be based solely on “Cloud vs. On-Prem”, but on privileges, data classes and operational responsibility. An external admin in the internal network is often riskier than a SaaS service with a clean tenancy model and clear evidence.
3) Core question in the selection process: Which operating model do you retain yourself — and what do you actually delegate?
Many later conflicts arise from unclear expectations: “Managed” does not automatically mean 24/7, does not automatically include patching, does not automatically include backup–RESTore tests. Therefore the operating model must be described in operational terms before contract conclusion.
Practical delineation: RACI and operational artifacts
RACI assigns roles: Responsible (executor), Accountable (owner), Consulted (consulted), Informed (to be informed). For audits it is particularly important that “Accountable” remains unambiguous, even if a service provider acts operationally.
Artifacts you should require mandatorily from critical service providers:
- Operations manual/runbooks (routine and emergency)
- Change process including approval rules and rollback
- Monitoring and alerting paths (including on-call arrangements)
- Backup and RESTore concept (including test evidence)
- Patch and vulnerability management (cycles, exceptions, risk approvals)
- Incident handling including evidence (preservation) and communication matrix
4) Compliance and evidence: Which documents really help?
Compliance is often mistaken for “paper”. What matters is whether evidence is verifiable and scope-appropriate. An ISO 27001 certificate can be valuable — but without a scope review it says little. A SOC 2 report (Type II) can be very helpful — but only if the controls match your risk profile and exceptions are understood and addressed.
Evidence that carries weight in practice
- ISMS evidence: ISO 27001 (scope, Statement of Applicability), internal policies, risk treatment.
- SOC 2 Type II: period, tested Trust Services Criteria, findings/exceptions, subservice organizations.
- Penetration test / vulnerability management evidence: frequency, scope, handling of findings (without you necessarily receiving full detailed reports).
- Data protection documents: AVV (data processing agreement), TOMs (technical and organizational measures), list of subprocessors, deletion and return concept.
Important: „We comply with GDPR“ is not a statement. For GDPR you need concrete building blocks: AVV, purpose, categories of data subjects, types of data, deletion/retention periods, technical measures, international transfers (e.g. Standard Contractual Clauses) and a robust subcontractor management.
5) Security criteria for selection: controls that matter in operation
For the selection decision you should formulate security criteria so they can later be operationalized. Instead of „high security“ you need verifiable requirements. Below are the control families that regularly prove decisive in supplier relationships.
Identities and privileged access (IAM/PAM)
IAM (Identity and Access Management) governs identities, roles and permissions. PAM (Privileged Access Management) controls especially powerful admin access, ideally time-limited and auditable.
- Individual users instead of shared accounts
- MFA (Multi-Factor Authentication) mandatory
- Just-in-Time / Just-Enough-Access where possible
- Access via jump hosts / bastion, no direct admin logins from the Internet
- Session recording or at least detailed audit logs for privileged actions
Audit question: Can you, in an incident, demonstrate who when what did, and can you revoke access within minutes?
Network and tenant separation
Segmentation is crucial especially for managed services: separate networks for management, production, backup, logging; clear firewall rules; and a documented exception policy. Tenant separation is central for SaaS providers: logical separation (tenant isolation), encryption and protection against data leakage due to misconfiguration.
Vulnerability and patch management
For selection and contract, the „patch cycle“ matters less than the ability to manage risk: How is prioritization handled (critical/high/medium), how are exceptions documented, how is compensation implemented (e.g. WAF rules, isolation), and how quickly can hotfixes be applied?
A simple but effective proof is a regular export from the ticket/vulnerability process (anonymized) showing: intake, assessment, deadline, implementation, review.
Logging, monitoring, forensic capability
This determines audit readiness: Without reliable logs (timestamps, integrity protection, retention) incidents remain a matter of opinion. It is particularly problematic for a service provider when logs reside with the provider but you, as the customer, carry the burden of proof.
- Defined log sources (auth, admin actions, system changes, API access)
- Centralized storage with an access concept (least privilege)
- Retention periods aligned with regulation and internal policies
- Integrity (tamper protection) and time synchronization (NTP)
6) Data protection and data sovereignty: AVV is only the beginning
Data protection is often reduced to the AVV (Data Processing Agreement) when selecting a service provider. That is risky because the critical questions lie in the technical implementation: Where is data processed, how is it encrypted, how is deletion performed, and how is data return handled during an exit?
Concrete requirements you should document
- Data location: regions/data centers, international transfers, legal bases.
- Encryption: in transit (TLS) and at REST; key management (KMS), access to keys.
- Backups: do they contain personal data? How are backups deleted? What retention applies?
- Data Subject Requests: support for access/deletion/export, deadlines and process.
- Role model: who is the controller, who is the processor, who is the subprocessor?
If you have strict requirements (e.g. keys under your own control, „Bring Your Own Key“), these requirements must be included in the must-have criteria before selection. Afterwards such points are often expensive or technically infeasible.
7) Subservice providers and supply chain: the blind spot in vendor management
Many risks do not arise with the selected partner, but in the chain: cloud hosting, 24/7 NOC, external development teams, support in other jurisdictions. Subservice providers are not inherently bad — but they must be transparent, controllable and contractually covered.
What you should insist on
- Current list of subservice providers with scope of services and data access
- Change process: prior notification and rights to object/terminate for critical changes
- „Flow-down“ clauses: security and data protection requirements apply throughout the chain
- Rights to review relevant evidence (e.g. SOC reports of the subservice organization)
8) Contract and SLA logic: security requirements must become measurable
Contracts rarely fail due to a missing paragraph; they fail because expectations are not measurable. SLAs (Service Level Agreements) describe service objectives (e.g. availability, response times). OLAs (Operational Level Agreements) are internal/operational agreements between teams or between service-provider units that make SLAs achievable in the first place.
Typical SLA components you should specify
- Incident classification: P1/P2/P3 definitions based on business impact
- Response and recovery times: not just „Response“, but „RESTore“
- Change windows: standard vs. emergency changes, documentation requirements
- Security SLAs: deadlines for remediation of critical vulnerabilities, patch cycles, exception rules
- Reporting: monthly service reports with defined metrics and deviation analysis
Important from an audit perspective: If you contractually define security controls, you also need a measurement and evidence routine. Otherwise a gap forms between contract and reality.
9) Audit perspective: what auditors typically want to see
Audits rarely check isolated technical details; they assess controllability: Is there a procedure, are decisions documented, are responsibilities clear, and are controls effective. For service-provider relationships many reviews boil down to three questions:
- Were risks assessed before engagement? (Due Diligence, tiering, approvals)
- Are controls operational and embedded in the contract? (SLA, security requirements, data protection, subcontractors)
- Can you demonstrate effectiveness? (reports, logs, tests, review records)
Audit-capable does not mean “document-heavy”. It means: make decisions traceable, document concisely, store evidence centrally, and perform regular reviews.
10) Practical checklist: due-diligence questions that actually help selection
The following list is intended as a pragmatic core. It does not replace an individual risk analysis, but it covers the typical decision criteria that later become relevant in operation and audit.
A) Scope and access
- Which systems/environments are in scope (Prod/Stage/Dev)?
- Which access types (VPN, jump host, API, ticket-only) are required?
- How are privileged accesses granted, logged and revoked?
B) Security Controls
- Which minimum standards apply (MFA, password policy, hardening, EDR/AV)?
- How is patch/vulnerability management performed, including prioritization and exceptions?
- How is logging implemented, who has access, what retention applies?
C) Data protection and data management
- AVV/TOMs: are they current and do they match the actual process?
- Data location, subcontractors, international transfers: are they described clearly?
- Deletion, return and backup retention: operationally achievable?
D) Operations, resilience, incident response
- Monitoring and on-call: schedules, escalation paths, communication channels?
- Backup/RESTore: are RESTores tested and documented?
- BCM/DR: are there tests, and what were the latest results?
E) Governance and evidence
- Which certifications/assurances (ISO/SOC) exist and what scope is covered?
- How are changes, incidents and reviews documented?
- Is there an audit right, and how is it implemented in practice (e.g. remote audit, report access)?
11) Scorecard approach: make the decision transparent (without pseudo-precision)
A scorecard helps make multiple providers comparable and justifies decisions to executive management, internal audit or data protection. Important: avoid pseudo-precision. Use few criteria, clear weighting, and document deviations with compensating measures.
Suggested weighting (example)
- 30% Security & access controls
- 25% Operations & resilience (RTO/RPO, runbooks, on-call)
- 20% Compliance & evidence (ISO/SOC/AVV, auditability)
- 15% Supply chain & subcontractor control
- 10% Commercial factors (cost model, transparency, flexibility)
For Tier-1 service providers there should be must-have criteria that may not be „down-weighted“ (e.g. MFA, individual accounts, AVV for personal data, exit arrangements, logging). If a must-have criterion is not met, the supplier is either excluded or a formally approved risk acceptance with compensation is required.
12) Exit and emergency planning: selection criterion, not an afterthought
Exit sounds like contract termination, but it is a security and operations lever: what happens in the event of insolvency, a severe incident, dispute, regulatory halt or strategic change? Without an exit plan the risk of lock-in increases, and in an emergency there is no time to hand over data and know-how cleanly.
Elements that belong in selection and contract
- Data return: formats, completeness, time windows, responsibilities
- Deletion confirmation: including backups/archive copies (as far as technically possible, described transparently)
- Delivery of documentation: runbooks, architecture, configurations, key/certificate inventory
- Transition support: defined hours/quotas, prioritized tasks, access for successor
- Emergency exit: special termination, access to systems/logs, freeze of changes
From an operational perspective it is particularly important that you receive not only „data“ but also operational capability: configurations, access models, monitoring setups and recovery paths.
13) Practical source blocks: templates for policies and audit commands
The following examples are intentionally generic and must be adapted to your environment. You can use them as a starting point for internal policies, supplier requirements or audit checks.
Policy template: minimum requirements for service provider access
POLICY: Third-party access to IT systems (minimum requirements)
1. Identities
- Only individual user accounts, no shared accounts.
- MFA mandatory for all external accesses.
- Role-based permissions (least privilege), time-limited admin rights (Just-in-Time).
2. Access paths
- Admin access exclusively via defined jump hosts/bastion or PAM solution.
- No direct administrative access from the Internet.
- Access only from approved networks/sources (Allowlist), where practicable.
3. Logging
- Authentication, privileged actions and configuration changes are logged centrally.
- Logs are protected against tampering and retained according to retention requirements.
4. Onboarding/Offboarding
- Approval process before provisioning, incl. ticket and owner.
- Deprovisioning within a defined timeframe after role change/project end.
5. Exceptions
- Exceptions only with documented risk assessment, expiration date and compensating controls.Audit commands (examples): traceability of admin accesses under Linux
For many environments it is important whether privileged actions are traceable. The following checks help verify typical fundamentals (auditd/journal/sudo). They are not a complete security assessment but a quick reality check during onboarding or review.
# Prüfen, ob sudo-Aktionen geloggt werden (Beispielpfade je nach Distribution)
sudo grep -R "^Defaults" /etc/sudoers /etc/sudoers.d 2>/dev/null | head
# Letzte sudo-Ereignisse (wenn über journald erfasst)
sudo journalctl -u sudo --since "7 days ago" 2>/dev/null | tail -n 50
# Prüfen, ob auditd aktiv ist
sudo systemctl status auditd --no-pager
# Audit-Regeln anzeigen (falls auditd genutzt wird)
sudo auditctl -l 2>/dev/null | head -n 50Important for evaluating a service provider is less which specific tool is used and more whether you achieve your objective: immutable, centralized, and analysable audit traces for relevant actions.
14) Assess costs and effort realistically: Security is not free, uncertainty is costlier
Risk-based selection saves money when it prevents you from buying “cheap” in the wrong place. Typical cost blocks that are underestimated in business cases:
- Transaction costs: onboarding, due diligence, contract review, tool integrations
- Control costs: reviews, reports, audits, pen tests, access recertification
- Incident costs: forensics, downtime, customer communication, regulatory notifications
- Lock-in costs: exit projects, data migration, knowledge transfer
A pragmatic approach: treat Tier-1 suppliers like internal critical systems. That does not mean “do everything yourself”, but establish controllability: clear roles, clear metrics, clear evidence.
15) Governance in daily operations: Who decides what – and who bears which responsibility?
Vendor governance only works if it is practiced in day-to-day operations. That requires established rituals and unambiguous responsibilities:
- Service Owner on the customer side: responsible for functional/operational matters, evaluates reports, manages the backlog
- Security/Compliance: defines mandatory controls, reviews deviations, approves exceptions
- Data Protection: assesses data flows, DPA/TOMs, transfers, deletion concepts
- Provider Manager (or Procurement/Legal): contract and escalation management
- IT Operations: integrates monitoring, logging, backup, changes, on-call processes
Quarterly Service Reviews (SLA, incidents, changes, security findings) and an annual Re-Assessment for Tier‑1/2 have proven effective. Not as a formality, but as a moment for hard questions: Have scope, subservice providers, data classes or access patterns changed?
Conclusion: Selecting IT service providers means purchasing controllable risks
A good selection decision follows a simple principle: The greater the access and the more critical the data or operational responsibility, the stricter the evidence, controls and exit rules must be. If you consistently apply tiering, scorecards and mandatory criteria, you avoid two extremes: excessive requirements for non‑critical suppliers and dangerous gaps for critical partners.
Rely on verifiable criteria (IAM/PAM, logging, patch processes, subservice-provider governance), anchor them contractually with measurable terms (SLA/Security SLAs) and plan exit and emergency handovers from the start. That turns vendor management from “gut feeling” into an auditable, operationally viable discipline.
For this topic, third-party risk management and supplier management are also important. The article places these aspects in a clear context and shows what matters in day-to-day operations.