[{"data":1,"prerenderedAt":2192},["ShallowReactive",2],{"guide-\u002Fguides\u002Fen\u002Fincident-management":3,"related-guides-en-statuspage-best-practices-uptime-monitoring-what-is-a-statuspage":617},{"id":4,"title":5,"body":6,"category":596,"description":597,"extension":598,"faq":599,"meta":600,"navigation":601,"path":602,"peerSlug":603,"publishedAt":604,"readingTime":605,"relatedComparisons":606,"relatedGuides":610,"seo":614,"stem":615,"updatedAt":599,"__hash__":616},"guides\u002Fguides\u002Fen\u002Fincident-management.md","Incident Management: Process, Roles and Communication",{"type":7,"value":8,"toc":571},"minimark",[9,14,18,21,24,29,32,55,58,62,65,69,78,82,85,89,92,96,107,111,114,118,121,124,206,209,212,257,269,273,276,307,310,314,317,320,325,329,332,338,344,347,373,384,388,391,416,419,422,426,429,449,456,460,463,476,480,483,501,516,520,523,561,565,568],[10,11,13],"h2",{"id":12},"what-is-incident-management","What Is Incident Management?",[15,16,17],"p",{},"Incident management is the structured process a team uses to detect an unplanned service disruption, coordinate its repair and learn from it afterward. An incident is an unplanned interruption or a noticeable degradation in the quality of a service — anything from a completely unreachable application to a checkout page that has slowed to a crawl. The primary goal of incident management is not to immediately find the deepest root cause, but to restore the service as quickly as possible.",[15,19,20],{},"This definition stays close to the ITIL understanding, where incident management is its own discipline within IT service management. You don't need to implement ITIL in detail to benefit from it. What matters is that everyone on the team means the same thing by \"an incident,\" knows who does what during a disruption, and that customer communication doesn't happen on the fly.",[15,22,23],{},"Incident management in plain terms: it's the emergency plan for your digital services. When something breaks, the process makes sure the right people come together quickly, the damage is contained, customers are informed honestly, and the team learns from the event afterward — instead of starting from scratch every time.",[25,26,28],"h3",{"id":27},"distinguishing-incident-problem-and-service-request","Distinguishing Incident, Problem and Service Request",[15,30,31],{},"Three terms often get blurred together in everyday work, but they mean different things:",[33,34,35,43,49],"ul",{},[36,37,38,42],"li",{},[39,40,41],"strong",{},"Incident"," is the active disruption. The focus is on restoration. Example: the API has been returning 500 errors for the past ten minutes.",[36,44,45,48],{},[39,46,47],{},"Problem"," is the underlying cause behind one or more incidents. The focus is on permanently eliminating the cause. Example: a memory leak that repeatedly takes the API down.",[36,50,51,54],{},[39,52,53],{},"Service request"," is a planned, standard request with no outage involved — such as the wish for a new user account or more storage. There is nothing to restore here.",[15,56,57],{},"This distinction is more than nitpicking. An incident is handled with high urgency, a problem with care and analytical depth, a service request through a calm, planned workflow. Throw everything into the same bucket and you either treat routine requests as emergencies or let real outages drown among ordinary tickets.",[10,59,61],{"id":60},"the-incident-lifecycle","The Incident Lifecycle",[15,63,64],{},"Every incident goes through the same phases, whether it's done in five minutes or only after hours of work. A clear lifecycle makes sure nobody skips steps, especially under pressure.",[25,66,68],{"id":67},"step-1-detection","Step 1: Detection",[15,70,71,72,77],{},"It all begins with detection. Ideally, your monitoring reports the disruption before the first customer notices it. The earlier the detection, the shorter the total downtime ends up being. How to monitor services reliably and which check types exist is covered in detail in the ",[73,74,76],"a",{"href":75},"\u002Fen\u002Fguides\u002Fuptime-monitoring","uptime monitoring"," guide. A disruption that only comes to light through a customer complaint has usually already done damage.",[25,79,81],{"id":80},"step-2-triage-and-classification","Step 2: Triage and Classification",[15,83,84],{},"Once an alert exists, triage follows. This is where the severity is assessed, the scope is estimated (Who is affected? Which functions?) and a decision is made about whether a formal incident process is even needed. Not every alert is a major incident. Triage prevents small things from triggering the whole machinery while real emergencies slip through.",[25,86,88],{"id":87},"step-3-response","Step 3: Response",[15,90,91],{},"In the response phase, the team forms up. Roles are assigned, the investigation begins, first hypotheses are tested. In parallel — and this is crucial — communication starts, both internally and externally. Customers would rather wait on an honest \"we're investigating\" message than hear nothing at all.",[25,93,95],{"id":94},"step-4-mitigation-and-resolution","Step 4: Mitigation and Resolution",[15,97,98,99,102,103,106],{},"There is an important distinction to make here. ",[39,100,101],{},"Mitigation"," means addressing the symptom and making the service available again — for instance through a rollback, a restart or by rerouting traffic. ",[39,104,105],{},"Resolution"," in the strict sense means permanently eliminating the actual cause. In practice, mitigation comes first: the service is running again, customers are taken care of. The clean root-cause fix often follows later, calmly, frequently as part of downstream problem management. Confusing these two steps leads teams to declare an incident \"solved\" too early.",[25,108,110],{"id":109},"step-5-follow-up-and-post-mortem","Step 5: Follow-Up and Post-Mortem",[15,112,113],{},"After restoration, the incident is technically over, but the process is not. The follow-up in the form of a post-mortem records what happened, why, and what needs to change. This step is the one most often skipped — and that is exactly why incidents repeat.",[10,115,117],{"id":116},"severity-and-status-levels","Severity and Status Levels",[15,119,120],{},"Severity and public status levels describe two different things: severity says internally how bad it is; the status level says externally where you currently stand in the handling process.",[15,122,123],{},"A common severity scale looks like this:",[125,126,127,146],"table",{},[128,129,130],"thead",{},[131,132,133,137,140,143],"tr",{},[134,135,136],"th",{},"Severity",[134,138,139],{},"Meaning",[134,141,142],{},"Example",[134,144,145],{},"Response",[147,148,149,164,178,192],"tbody",{},[131,150,151,155,158,161],{},[152,153,154],"td",{},"SEV1",[152,156,157],{},"Critical, complete or broad outage",[152,159,160],{},"Platform entirely unreachable",[152,162,163],{},"Immediate, all-hands",[131,165,166,169,172,175],{},[152,167,168],{},"SEV2",[152,170,171],{},"Significant degradation",[152,173,174],{},"Some users or core functions affected",[152,176,177],{},"Urgent, dedicated team",[131,179,180,183,186,189],{},[152,181,182],{},"SEV3",[152,184,185],{},"Minor degradation, workaround available",[152,187,188],{},"A single non-critical feature impaired",[152,190,191],{},"Soon, regular workflow",[131,193,194,197,200,203],{},[152,195,196],{},"SEV4\u002FSEV5",[152,198,199],{},"Cosmetic or trivial",[152,201,202],{},"Typo, slight rendering glitch",[152,204,205],{},"In the backlog, by priority",[15,207,208],{},"Important: severity is not the same as priority. Severity describes the degree of impact, priority the order of handling. A SEV2 incident can take a higher priority than another SEV2 depending on business context — for example when an important customer is affected.",[15,210,211],{},"The public status levels of a status page follow their own well-established scheme:",[125,213,214,223],{},[128,215,216],{},[131,217,218,221],{},[134,219,220],{},"Status",[134,222,139],{},[147,224,225,233,241,249],{},[131,226,227,230],{},[152,228,229],{},"Investigating",[152,231,232],{},"Disruption detected, cause still unclear, investigation underway",[131,234,235,238],{},[152,236,237],{},"Identified",[152,239,240],{},"Cause found, fix being applied",[131,242,243,246],{},[152,244,245],{},"Monitoring",[152,247,248],{},"Fix rolled out, team watching for stability",[131,250,251,254],{},[152,252,253],{},"Resolved",[152,255,256],{},"Service fully restored",[15,258,259,260,263,264,268],{},"On top of this comes ",[39,261,262],{},"Maintenance"," as its own type — a planned maintenance window is not an outage and should not be communicated as an incident. How to drive these levels cleanly to the outside is explored further in the ",[73,265,267],{"href":266},"\u002Fen\u002Fguides\u002Fstatuspage-best-practices","status page best practices"," guide.",[10,270,272],{"id":271},"roles-during-an-incident","Roles During an Incident",[15,274,275],{},"In a larger incident, it isn't enough for \"everyone to help somehow.\" Clear roles prevent chaos and duplicated effort. For small disruptions, one person can hold several roles; in a SEV1, they should be separate.",[33,277,278,289,295,301],{},[36,279,280,283,284,288],{},[39,281,282],{},"Incident Commander (IC):"," coordinates the entire incident, makes decisions and delegates. The IC deliberately does ",[285,286,287],"em",{},"not"," debug — they keep the overview instead of getting lost in the technical detail. Their job is steering, not troubleshooting.",[36,290,291,294],{},[39,292,293],{},"Communications \u002F Comms Lead:"," owns all communication, internally to stakeholders and externally to customers, including the status page updates. This keeps the technical crew free to work undisturbed.",[36,296,297,300],{},[39,298,299],{},"Operations \u002F Tech Lead:"," runs the technical investigation and implements the fix. This role is deep in the code, the logs and the infrastructure.",[36,302,303,306],{},[39,304,305],{},"Scribe:"," records the timeline — who did or noticed what, and when. This log later becomes the basis of the post-mortem and, along the way, takes a significant load off everyone's memory.",[15,308,309],{},"What an incident commander actually does can be summed up in one sentence: they are the calm point that keeps the process moving, drives decisions and makes sure every role knows its task — without chasing the disruption at the keyboard themselves.",[10,311,313],{"id":312},"escalation-on-call-and-alerting","Escalation, On-Call and Alerting",[15,315,316],{},"For an incident to reach anyone at all, you need an on-call duty and an escalation chain. On-call means that at any given time a defined person is responsible for incoming alerts — usually in a rotation, say weekly, so the load is shared fairly and nobody has to be reachable permanently.",[15,318,319],{},"The escalation chain governs what happens when the primary doesn't respond: primary → secondary → manager. If the first level doesn't react within a set deadline, the alert automatically moves to the next. This way no critical alert is left lying around just because someone happens not to pick up the phone.",[15,321,322,323,268],{},"The biggest enemy of a working on-call setup is alert fatigue. When too many or too imprecise alerts arrive, the on-call responders go numb and eventually miss the one alert that matters. Good alert thresholds, bundling related signals and protection against false alarms are therefore not niceties but a prerequisite for reliability. Which monitoring strategies reduce false alarms is described in the ",[73,324,76],{"href":75},[10,326,328],{"id":327},"incident-communication-internal-and-external","Incident Communication: Internal and External",[15,330,331],{},"Communication determines how an incident is perceived — often more strongly than the technical severity itself. There are two levels to separate here.",[15,333,334,337],{},[39,335,336],{},"Internally",", it's about coordination: who is working on what, which hypothesis is currently being tested, which stakeholders need to know? This communication is allowed to be technical and detailed.",[15,339,340,343],{},[39,341,342],{},"Externally",", toward customers, a different rule applies: speak the language of users, not of engineers. Nobody outside the team cares about the exact stack trace. Customers want to know: what isn't working, does it affect me, and when will I hear from you again?",[15,345,346],{},"A good external update contains four things:",[33,348,349,355,361,367],{},[36,350,351,354],{},[39,352,353],{},"Impact:"," what is concretely limited, in words users understand.",[36,356,357,360],{},[39,358,359],{},"Affected services:"," which parts of the offering are touched, which run normally.",[36,362,363,366],{},[39,364,365],{},"Current status:"," what you are doing right now, without false promises.",[36,368,369,372],{},[39,370,371],{},"Next update time:"," when the next message will come — even if there's nothing new by then.",[15,374,375,376,380,381,383],{},"On cadence: better to update regularly, even without new findings, than to go silent. A \"we're still investigating, next update in 30 minutes\" signals that someone is on it. Silence reads as loss of control. Your own status page is the right channel for this, because it takes the load off support and offers a single reliable source — the fundamentals are explained in the ",[73,377,379],{"href":378},"\u002Fen\u002Fguides\u002Fwhat-is-a-statuspage","what is a status page"," guide. How to shape tone and update rhythm in concrete terms is covered further in the ",[73,382,267],{"href":266},".",[10,385,387],{"id":386},"post-mortem-and-blameless-retrospective","Post-Mortem and Blameless Retrospective",[15,389,390],{},"The post-mortem is the step that turns an event into genuine learning. Without it, a team keeps fixing the same class of disruption over and over without ever finding the lever. A useful post-mortem is structured and records at least four components:",[33,392,393,399,404,410],{},[36,394,395,398],{},[39,396,397],{},"Timeline:"," what happened when, from detection to restoration. This is where the scribe's log pays off.",[36,400,401,403],{},[39,402,353],{}," who was affected, for how long and how severely — number of users, duration, affected functions.",[36,405,406,409],{},[39,407,408],{},"Root cause:"," the actual cause, not just the surface symptom.",[36,411,412,415],{},[39,413,414],{},"Action items:"," concrete measures, each with an owner and a due date. A post-mortem without named owners and deadlines is a wish list, not an improvement.",[15,417,418],{},"What matters is the blameless attitude. \"Blameless\" means the focus is on system and process failures, not on assigning blame to individuals. The assumption is that people acted reasonably with the knowledge available to them — if something still went wrong, the system was fragile. A blameless culture encourages open analysis: only those who fear no punishment will honestly tell what really happened. That honesty is exactly what you need to find the true cause.",[15,420,421],{},"In practice: schedule the post-mortem promptly, typically within a few days of a major incident, while the details are still fresh.",[10,423,425],{"id":424},"metrics-mtta-mttr-and-mtbf","Metrics: MTTA, MTTR and MTBF",[15,427,428],{},"Metrics make the quality of incident management measurable and show over time whether improvements are working. Three are central:",[33,430,431,437,443],{},[36,432,433,436],{},[39,434,435],{},"MTTA (Mean Time To Acknowledge):"," the average time from a triggered alert to its acknowledgment, that is, until someone begins handling it. A high MTTA points to gaps in on-call or to alert fatigue.",[36,438,439,442],{},[39,440,441],{},"MTTR (Mean Time To Recovery \u002F Repair \u002F Resolve):"," the average time from detection to restoration of the service. Probably the most-cited incident metric — it measures how quickly you are back online.",[36,444,445,448],{},[39,446,447],{},"MTBF (Mean Time Between Failures):"," total operating time divided by the number of failures. Here a higher value is better, because it means longer stretches of disruption-free operation.",[15,450,451,452,268],{},"MTTA and MTTR target response and repair, MTBF targets fundamental stability. Together they give an honest picture: how reliable is the service, and how well does the team respond when something does happen. How these metrics relate to plain availability in percent is shown in the ",[73,453,455],{"href":454},"\u002Fen\u002Fguides\u002Fuptime-percentage-explained","uptime percentage explained",[10,457,459],{"id":458},"from-monitoring-alert-to-incident","From Monitoring Alert to Incident",[15,461,462],{},"In many setups there is an invisible gap between the monitoring alert and the public incident. Monitoring fires, someone sees the alert — and then has to manually open a separate incident or status page tool, create an event and type out the text. This context switch costs time, exactly when time is scarcest.",[15,464,465,466,470,471,475],{},"The effect is measurable: every manual handoff step lengthens MTTA and MTTR. It also creates friction — under stress, the status page update is easily forgotten or posted too late, because it's an extra, deliberate action. Specialized incident tools like ",[73,467,469],{"href":468},"\u002Fen\u002Fincident-io-alternative","incident.io"," or ",[73,472,474],{"href":473},"\u002Fen\u002Filert-alternative","iLert"," handle coordination and escalation well, but bring no monitoring of their own. So you still need a separate monitoring tool and have to connect both worlds. The closer detection and incident workflow sit together, the smaller this gap becomes.",[10,477,479],{"id":478},"incident-management-with-livck","Incident Management with LIVCK",[15,481,482],{},"LIVCK is designed as one solution that unites monitoring and status page in a single tool — and thereby avoids exactly the context switch described above. The detected outage and the public incident live in the same system, instead of being manually reconciled across two tools.",[15,484,485,486,489,490,493,494,497,498,500],{},"For the actual incident handling, LIVCK offers a status workflow that follows the established scheme from Investigating through Identified and Monitoring to Resolved. ",[39,487,488],{},"Outage Linking"," groups several affected services with a shared cause into ",[285,491,492],{},"one"," incident — customers don't see five confusing separate notices, but one clear event. Alongside incidents, ",[39,495,496],{},"Announcements"," are available for general notices and ",[39,499,262],{}," for scheduled maintenance windows, so planned work stays cleanly separated from real disruptions.",[15,502,503,504,510,511,383],{},"Communication runs automatically: subscribers are notified across all relevant channels — email, Slack, Microsoft Teams, Telegram, Discord and webhooks. Against false alarms, majority detection protects you: an incident is only triggered after multiple independent probe locations have confirmed the disruption, which noticeably dampens alert fatigue in on-call. Escalation policies with acknowledge and postmortems attached directly to the incident round out the workflow. If you prefer self-hosting, you can install LIVCK in minutes via ",[73,505,509],{"href":506,"rel":507},"https:\u002F\u002Fhelp.livck.com",[508],"nofollow","Docker Compose","; an overview of the status page features is available at ",[73,512,515],{"href":513,"rel":514},"https:\u002F\u002Flivck.cloud\u002Fen",[508],"livck.cloud",[10,517,519],{"id":518},"common-mistakes","Common Mistakes",[15,521,522],{},"A few patterns keep recurring and cost time and nerves in every incident anew:",[33,524,525,531,537,543,549,555],{},[36,526,527,530],{},[39,528,529],{},"No defined incident commander."," Without a clear owner, everyone talks over each other and nobody makes decisions.",[36,532,533,536],{},[39,534,535],{},"No severity definition."," When it's unclear what separates a SEV1 from a SEV3, the team either over- or under-reacts.",[36,538,539,542],{},[39,540,541],{},"Customer communication too late or too rare."," Going silent during an outage damages trust more than the outage itself.",[36,544,545,548],{},[39,546,547],{},"Post-mortem with blame."," As soon as blame is in the room, the honest answers go quiet — and with them the chance to find the true cause.",[36,550,551,554],{},[39,552,553],{},"Alert fatigue."," Too many imprecise alerts cause the one important one to be missed.",[36,556,557,560],{},[39,558,559],{},"Separate, uncoupled tools."," Monitoring and incident communication in two worlds create a context switch that lengthens MTTA and MTTR.",[10,562,564],{"id":563},"conclusion","Conclusion",[15,566,567],{},"Incident management is not a tool you buy, but a process you practice. The building blocks are manageable: a clear lifecycle from detection to post-mortem, defined severity levels, named roles with an incident commander at the top, a reliable on-call with an escalation chain, and honest, regular communication in the language of users. Metrics like MTTA, MTTR and MTBF make it visible whether the work is having an effect, and a blameless post-mortem turns every event into learning rather than blame.",[15,569,570],{},"The biggest lever often sits where detection and response meet: the shorter the path from monitoring alert to public incident, the faster you are able to act again and the calmer you can communicate. Bringing monitoring and status page together in one system closes exactly this gap — and wins you, in a real emergency, the minutes that count.",{"title":572,"searchDepth":573,"depth":573,"links":574},"",2,[575,579,586,587,588,589,590,591,592,593,594,595],{"id":12,"depth":573,"text":13,"children":576},[577],{"id":27,"depth":578,"text":28},3,{"id":60,"depth":573,"text":61,"children":580},[581,582,583,584,585],{"id":67,"depth":578,"text":68},{"id":80,"depth":578,"text":81},{"id":87,"depth":578,"text":88},{"id":94,"depth":578,"text":95},{"id":109,"depth":578,"text":110},{"id":116,"depth":573,"text":117},{"id":271,"depth":573,"text":272},{"id":312,"depth":573,"text":313},{"id":327,"depth":573,"text":328},{"id":386,"depth":573,"text":387},{"id":424,"depth":573,"text":425},{"id":458,"depth":573,"text":459},{"id":478,"depth":573,"text":479},{"id":518,"depth":573,"text":519},{"id":563,"depth":573,"text":564},"Incident Management","Incident management explained: lifecycle, severity levels, roles, escalation, post-mortems and metrics like MTTR — plus professional incident communication.","md",null,{},true,"\u002Fguides\u002Fen\u002Fincident-management","incident-management","2026-06-24",13,[607,608,609],"incident-io","ilert","atlassian",[611,612,613],"statuspage-best-practices","uptime-monitoring","what-is-a-statuspage",{"title":5,"description":597},"guides\u002Fen\u002Fincident-management","2VH4r4fYEBtfLqksrXKwlX3QY11OJkxxm7sMfUK4qy0",[618,1157,1792],{"id":619,"title":620,"body":621,"category":1143,"description":1144,"extension":598,"faq":599,"meta":1145,"navigation":601,"path":1146,"peerSlug":611,"publishedAt":1147,"readingTime":1148,"relatedComparisons":1149,"relatedGuides":1152,"seo":1154,"stem":1155,"updatedAt":599,"__hash__":1156},"guides\u002Fguides\u002Fen\u002Fstatuspage-best-practices.md","Statuspage Best Practices — Transparency, Design & Incident Communication",{"type":7,"value":622,"toc":1086},[623,627,630,633,636,639,643,646,650,653,657,660,664,667,671,674,678,683,688,694,700,706,712,717,723,729,732,735,739,743,746,766,770,773,777,780,801,804,808,811,815,818,821,825,828,832,835,839,842,846,849,853,861,865,868,871,875,878,882,885,888,891,894,898,901,905,908,912,915,919,922,926,929,933,936,940,960,963,967,971,974,978,981,985,988,992,996,999,1003,1006,1010,1013,1017,1020,1024,1027,1031,1034,1038,1041,1045,1048,1052,1055,1057,1060,1080,1083],[10,624,626],{"id":625},"why-a-well-run-statuspage-matters","Why a Well-Run Statuspage Matters",[15,628,629],{},"A statuspage is more than a technical information display. It is the public face of your infrastructure and a direct communication channel between your engineering team and every user who depends on your service. In a world where downtime is inevitable, the quality of your statuspage determines whether users build trust or lose it.",[15,631,632],{},"The impact of a professionally maintained statuspage is measurable. Support teams consistently report a 30 to 50 percent reduction in incoming tickets during incidents when a statuspage is actively updated. Users who can inform themselves do not open tickets. They check the statuspage, see the current state, and wait.",[15,634,635],{},"Beyond ticket deflection, a statuspage is an instrument of customer retention. Companies that handle outages openly are perceived as more trustworthy than those that stay silent. Transparency is not a sign of weakness. It is a sign of professional maturity.",[15,637,638],{},"This guide covers the essential best practices that transform a statuspage from a forgotten subpage into a strategic communication tool.",[10,640,642],{"id":641},"transparent-communication-as-a-core-principle","Transparent Communication as a Core Principle",[15,644,645],{},"The most important rule for any statuspage: honesty over perfection. Users forgive outages. What they do not forgive is the feeling of being left in the dark.",[25,647,649],{"id":648},"proactive-over-reactive","Proactive Over Reactive",[15,651,652],{},"An incident should appear on the statuspage before the first support tickets arrive. If you wait until users report the problem, you have already lost the communication advantage. The statuspage is then no longer perceived as an information source but as a belated confirmation of what users already know.",[25,654,656],{"id":655},"plain-language-over-technical-jargon","Plain Language Over Technical Jargon",[15,658,659],{},"Statuspage updates must be understandable for all audiences. Not every reader is an engineer. An update like \"We are investigating increased error rates on the login service\" is better than \"HTTP 503 on auth-service-prod-3 after OOM kill in k8s cluster.\" Technical details belong in the post-mortem, not in the live update.",[25,661,663],{"id":662},"no-finger-pointing","No Finger-Pointing",[15,665,666],{},"A statuspage update describes the problem and the progress. It does not blame third-party providers, individual teams, or specific people. The phrasing \"Our cloud provider has reported a network issue affecting our service\" is factual and professional.",[10,668,670],{"id":669},"using-status-levels-effectively","Using Status Levels Effectively",[15,672,673],{},"A well-designed incident workflow gives every outage a clear structure. Instead of toggling between \"down\" and \"resolved,\" differentiated status levels allow precise communication of progress.",[25,675,677],{"id":676},"the-5-stage-workflow-in-detail","The 5-Stage Workflow in Detail",[15,679,680,682],{},[39,681,229],{}," — The starting point. A problem has been detected, the cause is unknown. A brief notice is sufficient: \"We are investigating reports of limited dashboard availability.\"",[15,684,685,687],{},[39,686,237],{}," — The root cause has been found. The team now knows what went wrong. \"The root cause has been identified: a faulty database migration is causing read timeouts.\"",[15,689,690,693],{},[39,691,692],{},"Acknowledged"," — The problem is recognized and prioritized, but the fix requires time. Particularly relevant for issues that cannot be resolved immediately.",[15,695,696,699],{},[39,697,698],{},"In Progress"," — Active work on the fix. Users see that something is happening. \"The team is performing a rollback of the database schema changes.\"",[15,701,702,705],{},[39,703,704],{},"Observing"," — The fix has been deployed, the team is monitoring. \"The fix has been rolled out. We are monitoring systems for the next 30 minutes.\"",[15,707,708,711],{},[39,709,710],{},"Under Review"," — Post-fix evaluation is underway. The incident is technically resolved, but the team is still checking for residual effects.",[15,713,714,716],{},[39,715,253],{}," — Everything is back to normal. Brief summary of what happened and what was done.",[15,718,719,722],{},[39,720,721],{},"Scheduled"," — For planned maintenance. Users are informed in advance.",[15,724,725,728],{},[39,726,727],{},"Closed"," — The incident is completed and archived.",[15,730,731],{},"The advantage of this model: users never have to guess where in the process the team currently is. Every status transition is a signal that active work is happening.",[15,733,734],{},"LIVCK implements this entire workflow natively. Every stage is built into the incident management system, including automatic notifications on status transitions.",[10,736,738],{"id":737},"incident-updates-frequency-tone-and-content","Incident Updates: Frequency, Tone, and Content",[25,740,742],{"id":741},"frequency","Frequency",[15,744,745],{},"During an active incident, the rule is: one update too many is better than one too few. A proven cadence:",[33,747,748,754,760],{},[36,749,750,753],{},[39,751,752],{},"First 30 minutes:"," An update every 10 to 15 minutes, even if it only says \"Still investigating.\"",[36,755,756,759],{},[39,757,758],{},"After 30 minutes:"," Every 20 to 30 minutes, provided the status is changing.",[36,761,762,765],{},[39,763,764],{},"For extended incidents:"," At least once per hour. Users who see no updates assume that nobody is working on the problem.",[25,767,769],{"id":768},"tone","Tone",[15,771,772],{},"Statuspage updates should be factual, calm, and solution-oriented. Neither panic nor minimization is appropriate. Phrases like \"minor issue\" or \"minimal impact\" are only acceptable when they match reality. Users who cannot access their service rarely perceive the impact as minimal.",[25,774,776],{"id":775},"content-of-a-good-update","Content of a Good Update",[15,778,779],{},"Every update should contain three elements:",[781,782,783,789,795],"ol",{},[36,784,785,788],{},[39,786,787],{},"Current state"," — What is happening right now?",[36,790,791,794],{},[39,792,793],{},"Next step"," — What is being done next?",[36,796,797,800],{},[39,798,799],{},"Expected timeline"," — When should users expect the next update?",[15,802,803],{},"An example: \"The login service remains degraded. The team has identified the cause as a corrupted cache entry and is currently performing a cache flush. We expect restoration within the next 15 minutes and will provide another update by 2:30 PM UTC at the latest.\"",[10,805,807],{"id":806},"informing-subscribers-the-right-way","Informing Subscribers the Right Way",[15,809,810],{},"A statuspage that nobody visits serves no purpose. A functioning notification system is therefore essential.",[25,812,814],{"id":813},"think-multi-channel","Think Multi-Channel",[15,816,817],{},"Not every user checks a website regularly. Professional statuspages offer notifications through multiple channels: email, Slack, Discord, Telegram, SMS. The choice of channel should be up to the user.",[15,819,820],{},"LIVCK natively supports email, Discord (webhook and bot), Slack, Telegram, SMS, and Pushover. All channels include throttling to prevent notification floods during major incidents.",[25,822,824],{"id":823},"newsletter-subscribers-vs-incident-subscribers","Newsletter Subscribers vs. Incident Subscribers",[15,826,827],{},"There is an important distinction between users who want to be notified about incidents and those who want general updates or announcements. A good statuspage supports both models and lets users decide which notifications they receive.",[25,829,831],{"id":830},"using-announcements","Using Announcements",[15,833,834],{},"Not every piece of communication is an incident. Planned migrations, new features, or infrastructure changes can be communicated through announcements without opening an incident. This keeps the incident history clean and gives the team a dedicated channel for proactive communication.",[10,836,838],{"id":837},"design-and-branding-projecting-professionalism","Design and Branding: Projecting Professionalism",[15,840,841],{},"The statuspage is often the first touchpoint during a crisis. The standards for design and branding must reflect this.",[25,843,845],{"id":844},"consistent-branding","Consistent Branding",[15,847,848],{},"The statuspage should visually belong to the main product. Same colors, same logo, same typography. A statuspage that looks like a foreign body undermines trust. Users wonder if they have landed on the right page.",[25,850,852],{"id":851},"custom-domain","Custom Domain",[15,854,855,856,860],{},"A statuspage at ",[857,858,859],"code",{},"status.yourproduct.com"," looks more professional than a generic subdomain of a third-party provider. Custom domains signal that the statuspage is an integral part of the product, not an afterthought.",[25,862,864],{"id":863},"clarity-over-creativity","Clarity Over Creativity",[15,866,867],{},"Statuspage design should prioritize information, not aesthetics. Large, clearly readable status indicators. Color coding for different states. No cluttered layouts. A user must be able to determine within three seconds whether everything is operational or whether there is a problem.",[15,869,870],{},"LIVCK offers full custom branding: logo, colors, custom CSS and custom domains are included in every plan.",[10,872,874],{"id":873},"monitoring-integration-why-monitoring-and-statuspage-belong-together","Monitoring Integration: Why Monitoring and Statuspage Belong Together",[15,876,877],{},"A statuspage without monitoring is reactive. The team learns about problems only when users report them. By the time the statuspage is updated, the incident has already caused damage.",[25,879,881],{"id":880},"automatic-detection-over-manual-reporting","Automatic Detection Over Manual Reporting",[15,883,884],{},"When monitoring and statuspage are integrated in one system, incidents can be triggered automatically as soon as a check fails. This reduces the mean time to notify dramatically.",[25,886,488],{"id":887},"outage-linking",[15,889,890],{},"An advanced concept is the automatic association of incidents with affected services. When the monitoring check for the API server fails, the corresponding service on the statuspage is automatically marked as affected. This outage linking eliminates manual steps and ensures the statuspage always reflects the actual system state.",[15,892,893],{},"LIVCK implements outage linking natively, connecting incidents with the services they affect without requiring manual intervention from the on-call engineer.",[25,895,897],{"id":896},"false-alarm-protection","False Alarm Protection",[15,899,900],{},"Not every failed check is a real outage. Network glitches, transient DNS issues, or brief timeouts can trigger false alarms. A good system filters these out before an incident is created. LIVCK therefore confirms an outage from multiple independent locations before an incident is created, preventing unnecessary noise on the statuspage and in notification channels.",[10,902,904],{"id":903},"communicating-maintenance-windows-professionally","Communicating Maintenance Windows Professionally",[15,906,907],{},"Planned maintenance is not an incident. It deserves its own communication strategy.",[25,909,911],{"id":910},"lead-time","Lead Time",[15,913,914],{},"Users should be notified at least 48 hours before scheduled maintenance. For larger changes with potential downtime, 5 to 7 days is appropriate. The announcement should clearly state the time window, the affected services, and the expected impact.",[25,916,918],{"id":917},"maintenance-windows-with-start-and-end-times","Maintenance Windows with Start and End Times",[15,920,921],{},"Vague statements like \"over the weekend\" are insufficient. Professional maintenance communication specifies exact times: \"Saturday, March 15, 2026, 02:00 to 04:00 UTC.\" Users can plan accordingly and inform their own stakeholders.",[25,923,925],{"id":924},"status-during-maintenance","Status During Maintenance",[15,927,928],{},"Even during planned maintenance, the statuspage should be updated. A status like \"Maintenance is proceeding as planned\" reassures users. If the maintenance takes longer than expected, an update becomes even more critical.",[10,930,932],{"id":931},"private-statuspages-for-internal-teams","Private Statuspages for Internal Teams",[15,934,935],{},"Not every statuspage is public. Internal statuspages for DevOps teams, management, or partners provide a protected communication channel with more technical detail.",[25,937,939],{"id":938},"use-cases","Use Cases",[33,941,942,948,954],{},[36,943,944,947],{},[39,945,946],{},"Internal infrastructure:"," Databases, message queues, internal APIs that end users never interact with directly.",[36,949,950,953],{},[39,951,952],{},"Partner integrations:"," Status overview for B2B partners who depend on your API.",[36,955,956,959],{},[39,957,958],{},"Regulated industries:"," Financial services, healthcare, or public sector organizations where certain information must not be publicly accessible.",[15,961,962],{},"Private pages with access control are included in every LIVCK plan. Teams can operate multiple statuspages for different audiences without needing separate tools.",[10,964,966],{"id":965},"uptime-metrics-and-sla-transparency","Uptime Metrics and SLA Transparency",[25,968,970],{"id":969},"uptime-calendar","Uptime Calendar",[15,972,973],{},"An uptime calendar shows historical availability at a glance. Users see not only the current status but also the reliability track record over past weeks and months. This transparency builds long-term trust and is often the first thing prospective customers evaluate.",[25,975,977],{"id":976},"sla-tracking","SLA Tracking",[15,979,980],{},"For B2B customers, service level agreements are not a formality but contractual obligations. A statuspage that openly displays SLA metrics demonstrates that the organization takes its own commitments seriously. It also provides a shared reference point for conversations between vendor and customer.",[25,982,984],{"id":983},"badges","Badges",[15,986,987],{},"Uptime badges embedded in documentation, repositories, or marketing pages are a simple but effective way to make availability visible. They serve as both a trust signal and a quick link to the statuspage. When availability is high, badges are a source of pride. When it drops, they are an incentive to improve.",[10,989,991],{"id":990},"common-mistakes-and-how-to-avoid-them","Common Mistakes and How to Avoid Them",[25,993,995],{"id":994},"mistake-1-only-updating-the-statuspage-during-major-outages","Mistake 1: Only Updating the Statuspage During Major Outages",[15,997,998],{},"Minor degradations belong on the statuspage too. If you only communicate during total outages, you train users not to trust the statuspage. When it always shows \"All Systems Operational\" despite users regularly experiencing issues, the page loses credibility.",[25,1000,1002],{"id":1001},"mistake-2-communicating-too-late","Mistake 2: Communicating Too Late",[15,1004,1005],{},"The most common mistake. The team wants to find the root cause before communicating. But users do not want to wait for a diagnosis. They want to know that the problem is acknowledged and someone is working on it. The first update can be as simple as: \"We are aware of issues affecting the dashboard and are investigating.\"",[25,1007,1009],{"id":1008},"mistake-3-no-post-mortem-after-major-incidents","Mistake 3: No Post-Mortem After Major Incidents",[15,1011,1012],{},"After a significant outage, users expect a thorough review. What happened? Why? What is being done to prevent recurrence? A post-mortem is not an admission of failure. It is proof of a functioning engineering process. The best post-mortems are blameless, specific, and include concrete action items.",[25,1014,1016],{"id":1015},"mistake-4-running-monitoring-and-statuspage-as-separate-systems","Mistake 4: Running Monitoring and Statuspage as Separate Systems",[15,1018,1019],{},"When the monitoring tool and the statuspage are not integrated, a manual process emerges: someone has to notice the outage, open the statuspage tool, and manually create an incident. This costs time and is error-prone. Integrated solutions like LIVCK eliminate this gap entirely.",[25,1021,1023],{"id":1022},"mistake-5-generic-design-with-no-brand-identity","Mistake 5: Generic Design With No Brand Identity",[15,1025,1026],{},"A statuspage that looks like a thousand others is not perceived as part of your product. Custom branding is not a luxury. It is a baseline requirement for professional external communication. Your statuspage should look like it was built by the same team that built your product.",[10,1028,1030],{"id":1029},"building-a-statuspage-culture","Building a Statuspage Culture",[15,1032,1033],{},"Technical implementation is only half the equation. The other half is organizational discipline.",[25,1035,1037],{"id":1036},"define-clear-ownership","Define Clear Ownership",[15,1039,1040],{},"Every team should know who is responsible for updating the statuspage during an incident. Ambiguity leads to delays. Whether it is the on-call engineer, a dedicated incident commander, or an automated system, the responsibility must be explicit.",[25,1042,1044],{"id":1043},"create-templates","Create Templates",[15,1046,1047],{},"Pre-written update templates for common scenarios accelerate response time. A template for \"Database degradation,\" \"Third-party provider outage,\" or \"Scheduled maintenance extended\" means the on-call engineer does not have to compose prose under pressure.",[25,1049,1051],{"id":1050},"review-and-iterate","Review and Iterate",[15,1053,1054],{},"After every major incident, review how the statuspage communication went. Was the first update timely? Were the status transitions accurate? Did subscribers receive notifications promptly? Each incident is an opportunity to improve the process.",[10,1056,564],{"id":563},[15,1058,1059],{},"A professional statuspage is not a side project. It is a strategic communication tool that builds trust, reduces support costs, and strengthens the relationship with users. The best practices distill into three core principles:",[781,1061,1062,1068,1074],{},[36,1063,1064,1067],{},[39,1065,1066],{},"Transparency:"," Communicate early, communicate honestly, communicate regularly.",[36,1069,1070,1073],{},[39,1071,1072],{},"Structure:"," Clear status levels, defined processes, consistent updates.",[36,1075,1076,1079],{},[39,1077,1078],{},"Integration:"," Monitoring and statuspage belong together. Manual processes lead to delays and errors.",[15,1081,1082],{},"LIVCK combines monitoring and statuspage in a single solution, with a full incident workflow, auto-detection, postmortems, multi-channel notifications, and full custom branding. Available as a self-hosted installation and as a cloud service, with all features included in every plan.",[15,1084,1085],{},"Organizations that commit to these best practices transform their statuspage from a technical obligation into a genuine competitive advantage.",{"title":572,"searchDepth":573,"depth":573,"links":1087},[1088,1089,1094,1097,1102,1107,1112,1117,1122,1125,1130,1137,1142],{"id":625,"depth":573,"text":626},{"id":641,"depth":573,"text":642,"children":1090},[1091,1092,1093],{"id":648,"depth":578,"text":649},{"id":655,"depth":578,"text":656},{"id":662,"depth":578,"text":663},{"id":669,"depth":573,"text":670,"children":1095},[1096],{"id":676,"depth":578,"text":677},{"id":737,"depth":573,"text":738,"children":1098},[1099,1100,1101],{"id":741,"depth":578,"text":742},{"id":768,"depth":578,"text":769},{"id":775,"depth":578,"text":776},{"id":806,"depth":573,"text":807,"children":1103},[1104,1105,1106],{"id":813,"depth":578,"text":814},{"id":823,"depth":578,"text":824},{"id":830,"depth":578,"text":831},{"id":837,"depth":573,"text":838,"children":1108},[1109,1110,1111],{"id":844,"depth":578,"text":845},{"id":851,"depth":578,"text":852},{"id":863,"depth":578,"text":864},{"id":873,"depth":573,"text":874,"children":1113},[1114,1115,1116],{"id":880,"depth":578,"text":881},{"id":887,"depth":578,"text":488},{"id":896,"depth":578,"text":897},{"id":903,"depth":573,"text":904,"children":1118},[1119,1120,1121],{"id":910,"depth":578,"text":911},{"id":917,"depth":578,"text":918},{"id":924,"depth":578,"text":925},{"id":931,"depth":573,"text":932,"children":1123},[1124],{"id":938,"depth":578,"text":939},{"id":965,"depth":573,"text":966,"children":1126},[1127,1128,1129],{"id":969,"depth":578,"text":970},{"id":976,"depth":578,"text":977},{"id":983,"depth":578,"text":984},{"id":990,"depth":573,"text":991,"children":1131},[1132,1133,1134,1135,1136],{"id":994,"depth":578,"text":995},{"id":1001,"depth":578,"text":1002},{"id":1008,"depth":578,"text":1009},{"id":1015,"depth":578,"text":1016},{"id":1022,"depth":578,"text":1023},{"id":1029,"depth":573,"text":1030,"children":1138},[1139,1140,1141],{"id":1036,"depth":578,"text":1037},{"id":1043,"depth":578,"text":1044},{"id":1050,"depth":578,"text":1051},{"id":563,"depth":573,"text":564},"Best Practices","The most important best practices for statuspages: transparent communication, proper incident management, professional design and monitoring integration.",{},"\u002Fguides\u002Fen\u002Fstatuspage-best-practices","2026-02-25",11,[609,1150,1151],"instatus","hyperping",[613,1153],"self-hosted-statuspage",{"title":620,"description":1144},"guides\u002Fen\u002Fstatuspage-best-practices","rKorgsy9A22r-xOwxWVSKLcU0DeNgJOfh9pt919u9Mo",{"id":1158,"title":1159,"body":1160,"category":245,"description":1780,"extension":598,"faq":599,"meta":1781,"navigation":601,"path":1782,"peerSlug":612,"publishedAt":604,"readingTime":605,"relatedComparisons":1783,"relatedGuides":1787,"seo":1789,"stem":1790,"updatedAt":599,"__hash__":1791},"guides\u002Fguides\u002Fen\u002Fuptime-monitoring.md","Uptime Monitoring: Fundamentals, Check Types and Choosing the Right Tool",{"type":7,"value":1161,"toc":1758},[1162,1166,1169,1172,1175,1189,1193,1196,1203,1207,1210,1213,1233,1239,1243,1246,1250,1253,1286,1290,1293,1297,1300,1304,1307,1311,1314,1318,1321,1324,1328,1331,1425,1429,1432,1435,1455,1458,1461,1465,1468,1488,1491,1498,1502,1505,1508,1514,1518,1521,1565,1583,1587,1608,1615,1618,1625,1670,1703,1707,1710,1736,1739,1741,1744,1754],[10,1163,1165],{"id":1164},"what-is-uptime-monitoring","What Is Uptime Monitoring?",[15,1167,1168],{},"Uptime monitoring is the continuous observation of a system's reachability from the outside. A monitoring service checks at fixed intervals whether your website, your API or your server responds — and whether the response is correct. That is the whole point of availability monitoring: not just \"is the server running\", but \"can a user actually reach and use the system\".",[15,1170,1171],{},"Technically, uptime monitoring is a black-box, or outside-in, form of observation. The monitoring system has no access to your source code or internal metrics. It looks at your system from the outside, exactly the way a real visitor would — over the network, with a request, evaluating the response. That is also its greatest strength: you measure precisely what your users experience, not just what an internal health check claims.",[15,1173,1174],{},"In website monitoring and server monitoring, two basic modes of operation complement each other:",[33,1176,1177,1183],{},[36,1178,1179,1182],{},[39,1180,1181],{},"Pull monitoring:"," The monitoring system actively queries the target. It sends an HTTP request, opens a TCP connection or sends a ping and evaluates the response. This is the standard case for publicly reachable services.",[36,1184,1185,1188],{},[39,1186,1187],{},"Push monitoring (heartbeat):"," The monitored system actively reports in to the monitoring service. If the expected signal fails to arrive, an alert fires. This suits cron jobs, background processes and systems behind a firewall that cannot be reached directly from the outside.",[25,1190,1192],{"id":1191},"difference-from-apm-and-performance-monitoring","Difference From APM and Performance Monitoring",[15,1194,1195],{},"Uptime monitoring is often confused with Application Performance Monitoring (APM), yet the two answer different questions. APM is white-box observation: it measures internal metrics such as database query times, error rates and code traces, and needs instrumentation directly in the application code.",[15,1197,1198,1199,1202],{},"The decisive difference can be put simply: ",[39,1200,1201],{},"Uptime monitoring tells you THAT something is broken. APM tells you WHY."," Uptime monitoring detects that the checkout page returns a 500. APM shows you that behind it sits a timeout to the payment database. The two complement each other, and neither replaces the other. For most websites, APIs and servers, solid uptime monitoring is the mandatory baseline — it is cheap, quick to set up and covers the most important question: is the system available to my users right now?",[10,1204,1206],{"id":1205},"why-monitoring-is-business-critical","Why Monitoring Is Business-Critical",[15,1208,1209],{},"The core principle: you learn about an outage first — not your customer. Without monitoring, the first report of an outage is often an annoyed user, a tweet or a support ticket. By that point the outage has already been running for a while, and you have lost control of the communication.",[15,1211,1212],{},"In practice, the cost of downtime falls into three areas:",[33,1214,1215,1221,1227],{},[36,1216,1217,1220],{},[39,1218,1219],{},"Revenue:"," Every outage in the ordering, booking or payment flow costs money directly. A shop that is unreachable for an hour loses not only that hour's orders, but often customers who never come back.",[36,1222,1223,1226],{},[39,1224,1225],{},"Trust and reputation:"," Repeated or uncommunicated outages damage trust — especially with B2B customers who depend on your service. Those who communicate transparently about incidents lose far less trust than those who stay silent.",[36,1228,1229,1232],{},[39,1230,1231],{},"SEO:"," Longer or frequent outages can affect crawling. If the search engine crawler repeatedly finds a page unreachable, that can affect indexing and ranking over the medium term.",[15,1234,1235,1236,383],{},"What an hour of downtime actually costs depends heavily on the business model and cannot be stated as a blanket figure. An honest estimate works if you take your average revenue per hour and add the indirect costs — support effort, reputational damage, lost new customers. Even this simple calculation usually shows quickly that a monitoring tool for a few euros a month is cheap insurance. How much downtime different availability targets even permit is something we work through in detail in the guide ",[73,1237,1238],{"href":454},"Uptime percentage explained",[10,1240,1242],{"id":1241},"the-most-important-check-types-in-detail","The Most Important Check Types in Detail",[15,1244,1245],{},"Not all uptime monitoring is the same. Depending on what you are monitoring, you need different check types. Anyone who only checks the home page over HTTP overlooks a whole range of failure scenarios.",[25,1247,1249],{"id":1248},"https-check","HTTP(S) Check",[15,1251,1252],{},"The HTTP(S) check is the most common type, and it does more than just verify that a page is reachable. It additionally evaluates:",[33,1254,1255,1277],{},[36,1256,1257,1260,1261,1264,1265,1268,1269,1272,1273,1276],{},[39,1258,1259],{},"The status code:"," A ",[857,1262,1263],{},"200 OK"," is good; a ",[857,1266,1267],{},"4xx"," (client error) or ",[857,1270,1271],{},"5xx"," (server error) signals a problem. A ",[857,1274,1275],{},"503 Service Unavailable"," is a classic outage signal.",[36,1278,1279,1282,1283,1285],{},[39,1280,1281],{},"Optionally a keyword or content in the body:"," This matters more than it sounds. A broken page can incorrectly return a ",[857,1284,1263],{}," and still show only a blank page or an error message. A content check looks for an expected term (for example \"Add to cart\") and raises an alert if it is missing — even when the status code is fine.",[25,1287,1289],{"id":1288},"tcpport-check","TCP\u002FPort Check",[15,1291,1292],{},"A TCP\u002Fport check verifies whether a specific port is open and accepting connections. This is relevant for services that do not run over HTTP: a database on port 5432 (PostgreSQL), a mail server on port 25 (SMTP) or any other network-based service. The check tells you whether the service accepts connections at all.",[25,1294,1296],{"id":1295},"icmpping-check","ICMP\u002FPing Check",[15,1298,1299],{},"An ICMP, or ping, check verifies the basic network reachability of a host. It tells you whether the server responds on the network — but not whether the services running on it work. A server can answer a ping while the web server has long since crashed. Ping is therefore a good first indicator, but no full substitute for a service-level check.",[25,1301,1303],{"id":1302},"ssl-certificate-check","SSL Certificate Check",[15,1305,1306],{},"Expired TLS certificates are one of the most common avoidable causes of outages. An SSL certificate check warns in good time before expiry — typically a few days ahead. That leaves enough time to renew the certificate before visitors see a security warning in their browser.",[25,1308,1310],{"id":1309},"dns-check","DNS Check",[15,1312,1313],{},"A DNS check verifies correct name resolution. If DNS fails or is misconfigured, your application is unreachable even though the server runs perfectly. DNS problems are insidious because they often surface only after configuration changes.",[25,1315,1317],{"id":1316},"heartbeatcron-monitoring","Heartbeat\u002FCron Monitoring",[15,1319,1320],{},"Heartbeat monitoring covers things that are neither an open port nor a URL: cron jobs, backup scripts, queue workers and other background processes. The principle is the reverse of pull monitoring: after a successful run, the job sends a ping to a URL. If the expected ping fails to arrive within the defined time window, the job is treated as failed and an alert fires.",[15,1322,1323],{},"This solves a problem that HTTP checks cannot address: a nightly backup script has no port you could check from the outside. If it does not complete, without a heartbeat nobody notices — until the backup is needed and missing.",[25,1325,1327],{"id":1326},"apimanual-check","API\u002FManual Check",[15,1329,1330],{},"Some states are known only to your own application. With an API, or manual, check you can set a status programmatically — for example to reflect a business-logic state that a generic check cannot detect.",[125,1332,1333,1346],{},[128,1334,1335],{},[131,1336,1337,1340,1343],{},[134,1338,1339],{},"Check type",[134,1341,1342],{},"Checks",[134,1344,1345],{},"Typical use",[147,1347,1348,1359,1370,1381,1392,1403,1414],{},[131,1349,1350,1353,1356],{},[152,1351,1352],{},"HTTP(S)",[152,1354,1355],{},"Status code + content",[152,1357,1358],{},"Websites, APIs",[131,1360,1361,1364,1367],{},[152,1362,1363],{},"TCP\u002FPort",[152,1365,1366],{},"Open port, connection",[152,1368,1369],{},"Databases, mail servers",[131,1371,1372,1375,1378],{},[152,1373,1374],{},"ICMP\u002FPing",[152,1376,1377],{},"Network reachability",[152,1379,1380],{},"Basic server check",[131,1382,1383,1386,1389],{},[152,1384,1385],{},"SSL",[152,1387,1388],{},"Certificate expiry",[152,1390,1391],{},"HTTPS services",[131,1393,1394,1397,1400],{},[152,1395,1396],{},"DNS",[152,1398,1399],{},"Name resolution",[152,1401,1402],{},"Domain configuration",[131,1404,1405,1408,1411],{},[152,1406,1407],{},"Heartbeat",[152,1409,1410],{},"Missing signal",[152,1412,1413],{},"Cron jobs, backups, workers",[131,1415,1416,1419,1422],{},[152,1417,1418],{},"Manual\u002FAPI",[152,1420,1421],{},"Custom status via API",[152,1423,1424],{},"Business logic",[10,1426,1428],{"id":1427},"check-interval-and-check-locations","Check Interval and Check Locations",[15,1430,1431],{},"The check interval determines how quickly you learn about an outage. With a 5-minute interval, an outage can go undetected for up to five minutes before the next check registers it. For many production systems that is too long.",[15,1433,1434],{},"As a rule of thumb:",[33,1436,1437,1443,1449],{},[36,1438,1439,1442],{},[39,1440,1441],{},"1 minute"," is the standard for production systems. A good compromise between fast detection and acceptable load.",[36,1444,1445,1448],{},[39,1446,1447],{},"30 seconds"," makes sense for critical systems where every minute counts — such as payment infrastructure or central APIs.",[36,1450,1451,1454],{},[39,1452,1453],{},"5 minutes or more"," is enough for non-critical secondary systems where short outages cause no harm.",[15,1456,1457],{},"A second, often underestimated factor is checking from multiple locations. If a target is checked from a single location only, a local network problem between the check point and the target can look like an outage — even though your application runs perfectly for everyone else. Multi-location monitoring checks from several geographic points and compares the results.",[15,1459,1460],{},"This is also where false-alarm protection comes in. Before an alert fires, the system should confirm a suspected failure: several check points must see the error (multi-location confirmation) and\u002For it is rechecked after a short delay (retry\u002Fconfirmation). Many transient glitches resolve themselves within 5 to 10 seconds — a brief recheck prevents them from turning into an alert.",[10,1462,1464],{"id":1463},"alerting-and-escalation","Alerting and Escalation",[15,1466,1467],{},"Monitoring that detects an outage but reaches no one is worthless. Alerting is therefore just as important as the check itself. In practice you need several channels, because different situations demand different levels of attention:",[33,1469,1470,1476,1482],{},[36,1471,1472,1475],{},[39,1473,1474],{},"Email"," for documentation and non-urgent notices.",[36,1477,1478,1481],{},[39,1479,1480],{},"Slack, Discord or Telegram"," for quick information in team chat.",[36,1483,1484,1487],{},[39,1485,1486],{},"SMS or Pushover"," for critical alerts that must arrive even at night.",[15,1489,1490],{},"The most important and most frequently underestimated aspect is alert fatigue. Anyone who receives too many or constantly repeated alerts becomes desensitized — and eventually misses the one alert that really matters. The remedies are throttling (not reporting every repetition again), grouping related alerts and clearly defined escalation paths: who is informed first, who takes over if no one responds?",[15,1492,1493,1494,268],{},"This is exactly where monitoring becomes a process. Anyone working with on-call rotations needs clear rules for when an alert escalates and to whom. How you get from the first notification to a structured response is something we cover in detail in the ",[73,1495,1497],{"href":1496},"\u002Fen\u002Fguides\u002Fincident-management","Incident management",[10,1499,1501],{"id":1500},"monitoring-and-statuspage-belong-together","Monitoring and Statuspage Belong Together",[15,1503,1504],{},"An alert informs your team. But what about your users? When an outage is confirmed, customers want to know what is going on — without having to open a support ticket first. This is exactly where the loop between monitoring and statuspage closes.",[15,1506,1507],{},"Ideally, a confirmed outage leads not only to an internal alert, but can be turned directly into a public statuspage update: a service is set to \"degraded\", subscribers are notified automatically, and your team can communicate the incident transparently while working on the fix. That takes pressure off support and builds trust.",[15,1509,1510,1511,383],{},"For this to work, monitoring and statuspage should not be two separate tools that you sync manually. What matters in a good statuspage — from service structure to incident communication — is something we have summarized in the ",[73,1512,1513],{"href":266},"Statuspage best practices",[10,1515,1517],{"id":1516},"criteria-for-choosing-a-tool","Criteria for Choosing a Tool",[15,1519,1520],{},"Which monitoring tool is the right one depends on your requirements. The following criteria help you compare offerings on the merits, rather than being dazzled by feature lists:",[33,1522,1523,1529,1535,1541,1547,1553,1559],{},[36,1524,1525,1528],{},[39,1526,1527],{},"Check types:"," Does the tool cover what you actually monitor? Plain HTTP is rarely enough — TCP, SSL and heartbeat belong in most setups.",[36,1530,1531,1534],{},[39,1532,1533],{},"Check interval:"," Does the plan you need offer an interval that is fast enough? Some cheap tiers are capped at 5 minutes.",[36,1536,1537,1540],{},[39,1538,1539],{},"Multi-location:"," Does it check from several locations to avoid false alarms?",[36,1542,1543,1546],{},[39,1544,1545],{},"Self-hosting option:"," Can you run the monitoring on your own infrastructure when you need full control?",[36,1548,1549,1552],{},[39,1550,1551],{},"Data privacy and server location:"," Where is the data processed? For companies with customers in the EU, a European or German server location is often mandatory.",[36,1554,1555,1558],{},[39,1556,1557],{},"Free tier:"," Is there a free entry point that lets you test the tool for real?",[36,1560,1561,1564],{},[39,1562,1563],{},"Incident integration:"," Can an alert be turned directly into statuspage communication, or does monitoring stay an isolated silo?",[15,1566,1567,1568,1572,1573,1577,1578,1582],{},"If you want to weigh specific providers against each other, you will find detailed comparisons — for example for ",[73,1569,1571],{"href":1570},"\u002Fen\u002Fuptimerobot-alternative","UptimeRobot",", ",[73,1574,1576],{"href":1575},"\u002Fen\u002Fbetter-stack-alternative","Better Stack"," and ",[73,1579,1581],{"href":1580},"\u002Fen\u002Fhetrixtools-alternative","HetrixTools",". There you can see in detail where each one's strengths and limits lie, particularly when it comes to server location and the link between monitoring and statuspage.",[10,1584,1586],{"id":1585},"uptime-monitoring-with-livck","Uptime Monitoring With LIVCK",[15,1588,1589,1590,1572,1592,1572,1594,1572,1597,1572,1599,1572,1601,1603,1604,1607],{},"LIVCK combines uptime monitoring and statuspage in a single tool — that is the essential difference from solutions that only monitor or only provide a statuspage. For monitoring, seven check types are available: ",[39,1591,1352],{},[39,1593,1363],{},[39,1595,1596],{},"Ping",[39,1598,1396],{},[39,1600,1385],{},[39,1602,1407],{}," (for cron jobs and background processes) and ",[39,1605,1606],{},"Manual"," via the API. This covers publicly reachable services, certificate expiry, DNS configuration and jobs behind a firewall.",[15,1609,1610,1611,1614],{},"Against false alarms, LIVCK works with ",[39,1612,1613],{},"majority detection across independent probe locations",": an incident is only triggered once the majority of locations confirm the error — checks run from up to eight locations, six of them in the DACH region. Transient glitches that resolve themselves after a few seconds therefore do not lead to unnecessary alerts and alert fatigue. Notifications run over email, SMS, voice call, Slack, Microsoft Teams, Telegram, Discord and webhooks.",[15,1616,1617],{},"The decisive advantage is the direct statuspage integration: a confirmed outage becomes a public update with automatic subscriber notification, without switching tools. Statuspage features such as custom branding, private pages and the uptime calendar are included in every plan, at no extra cost.",[15,1619,1620,1621,1624],{},"For operation, you have a choice. ",[39,1622,1623],{},"Self-hosted via Docker Compose"," is set up in minutes:",[1626,1627,1631],"pre",{"className":1628,"code":1629,"language":1630,"meta":572,"style":572},"language-bash shiki shiki-themes github-light github-dark","curl -fsSL https:\u002F\u002Fget.livck.com -o docker-compose.yml\ndocker compose up -d\n","bash",[857,1632,1633,1656],{"__ignoreMap":572},[1634,1635,1638,1642,1646,1650,1653],"span",{"class":1636,"line":1637},"line",1,[1634,1639,1641],{"class":1640},"sScJk","curl",[1634,1643,1645],{"class":1644},"sj4cs"," -fsSL",[1634,1647,1649],{"class":1648},"sZZnC"," https:\u002F\u002Fget.livck.com",[1634,1651,1652],{"class":1644}," -o",[1634,1654,1655],{"class":1648}," docker-compose.yml\n",[1634,1657,1658,1661,1664,1667],{"class":1636,"line":573},[1634,1659,1660],{"class":1640},"docker",[1634,1662,1663],{"class":1648}," compose",[1634,1665,1666],{"class":1648}," up",[1634,1668,1669],{"class":1644}," -d\n",[15,1671,1672,1673,1678,1679,1682,1683,1686,1687,1692,1693,1697,1698,1702],{},"A modest entry-level server is already enough, such as the ",[73,1674,1677],{"href":1675,"rel":1676},"https:\u002F\u002Flivck.com\u002Fr\u002Fhetzner",[508],"Hetzner Cloud CX22 from about €4.75\u002Fmonth"," — the only important point is that the monitoring does not run on the same infrastructure as the systems it monitors. Alternatively, a ",[39,1680,1681],{},"managed service on German servers"," is available, and the ",[39,1684,1685],{},"LIVCK Cloud Free plan"," (€0, 20 services, 120s interval, 1 statuspage) is in public beta — access via the waitlist. Because LIVCK comes from Germany and works GDPR by design, it is particularly well suited to regulated industries. You can find more on scope and installation in the ",[73,1688,1691],{"href":1689,"rel":1690},"https:\u002F\u002Fdocs.livck.cloud",[508],"LIVCK documentation"," and on the ",[73,1694,1696],{"href":513,"rel":1695},[508],"statuspage overview",". If you are generally interested in running things yourself, the guide on the ",[73,1699,1701],{"href":1700},"\u002Fen\u002Fguides\u002Fself-hosted-statuspage","self-hosted statuspage"," is worth a look too.",[10,1704,1706],{"id":1705},"common-monitoring-mistakes","Common Monitoring Mistakes",[15,1708,1709],{},"Even with the right tool, there are typical mistakes that undermine monitoring:",[33,1711,1712,1718,1724,1730],{},[36,1713,1714,1717],{},[39,1715,1716],{},"Only checking the home page:"," The home page can run perfectly while login, checkout or API have failed. Monitor the business-critical paths, not just the shop window.",[36,1719,1720,1723],{},[39,1721,1722],{},"Intervals that are too long:"," A 15-minute interval may save a few cents, but it leaves an outage undetected for a quarter of an hour. For critical systems that is too slow.",[36,1725,1726,1729],{},[39,1727,1728],{},"Monitoring on the same infrastructure:"," If monitoring runs on the same server or in the same data center as the monitored systems, it fails together with them — exactly when you need it most. Monitoring always belongs on independent infrastructure.",[36,1731,1732,1735],{},[39,1733,1734],{},"No alerting escalation path:"," An alert that only lands in an ignored email inbox helps no one. Define who is notified when, and what happens if no one responds.",[15,1737,1738],{},"These four points cost nothing but a little attention during setup — and in practice they decide whether your monitoring holds up when it matters or not.",[10,1740,564],{"id":563},[15,1742,1743],{},"Uptime monitoring is the cheapest insurance you can take out for your digital services: you learn about an outage before your customers do. What matters is that the monitoring fits what you run — with the right check types (not just HTTP, but also TCP, SSL and heartbeat for cron jobs), a sufficiently short check interval, false-alarm protection through multiple confirmation and a well-thought-out alerting path.",[15,1745,1746,1747,1749,1750,383],{},"You get the most out of it when monitoring and statuspage do not sit in separate silos but work together: a confirmed outage turns directly into transparent communication to your users. How quickly you need to react depends on your availability target — you will find the concrete numbers in the guide ",[73,1748,1238],{"href":454},". Finally, make sure to run the monitoring on independent infrastructure and — if data privacy matters to you — to check server location and ",[73,1751,1753],{"href":1752},"\u002Fen\u002Fguides\u002Fgdpr-statuspage","GDPR compliance",[1755,1756,1757],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":572,"searchDepth":573,"depth":573,"links":1759},[1760,1763,1764,1773,1774,1775,1776,1777,1778,1779],{"id":1164,"depth":573,"text":1165,"children":1761},[1762],{"id":1191,"depth":578,"text":1192},{"id":1205,"depth":573,"text":1206},{"id":1241,"depth":573,"text":1242,"children":1765},[1766,1767,1768,1769,1770,1771,1772],{"id":1248,"depth":578,"text":1249},{"id":1288,"depth":578,"text":1289},{"id":1295,"depth":578,"text":1296},{"id":1302,"depth":578,"text":1303},{"id":1309,"depth":578,"text":1310},{"id":1316,"depth":578,"text":1317},{"id":1326,"depth":578,"text":1327},{"id":1427,"depth":573,"text":1428},{"id":1463,"depth":573,"text":1464},{"id":1500,"depth":573,"text":1501},{"id":1516,"depth":573,"text":1517},{"id":1585,"depth":573,"text":1586},{"id":1705,"depth":573,"text":1706},{"id":563,"depth":573,"text":564},"Uptime monitoring explained: check types (HTTP, TCP, heartbeat), intervals, alerting and how to choose a tool — including self-hosting and GDPR.",{},"\u002Fguides\u002Fen\u002Fuptime-monitoring",[1784,1785,1786],"uptimerobot","better-stack","hetrixtools",[1788,603,611],"uptime-percentage-explained",{"title":1159,"description":1780},"guides\u002Fen\u002Fuptime-monitoring","QJCLMWfRd3cgXQdDpVJ1EpANeP0-nDxpFh2lgy1k3b0",{"id":1793,"title":1794,"body":1795,"category":2173,"description":2174,"extension":598,"faq":2175,"meta":2183,"navigation":601,"path":2184,"peerSlug":2185,"publishedAt":1147,"readingTime":2186,"relatedComparisons":2187,"relatedGuides":2188,"seo":2189,"stem":2190,"updatedAt":599,"__hash__":2191},"guides\u002Fguides\u002Fen\u002Fwhat-is-a-statuspage.md","What Is a Statuspage? — Definition, Benefits & Best Practices",{"type":7,"value":1796,"toc":2128},[1797,1800,1803,1809,1812,1815,1819,1823,1826,1830,1833,1837,1840,1844,1847,1851,1854,1858,1861,1865,1868,1872,1875,1879,1882,1886,1889,1893,1896,1900,1903,1907,1911,1914,1917,1921,1924,1928,1931,1935,1938,1942,1945,1949,1952,1956,1959,1963,1966,1970,1973,1977,1980,1984,1987,1991,1994,1998,2001,2004,2008,2012,2015,2019,2022,2026,2029,2033,2036,2040,2043,2047,2050,2056,2062,2068,2074,2080,2084,2087,2093,2099,2105,2111,2117,2119,2122,2125],[10,1798,1799],{"id":613},"What Is a Statuspage?",[15,1801,1802],{},"A statuspage is a publicly accessible web page that displays the real-time operational status of a company's services and systems. It serves as a single source of truth for customers, partners, and internal teams to check availability, view ongoing incidents, review scheduled maintenance, and assess historical uptime data.",[15,1804,1805,1806],{},"At its core, a statuspage answers one question: ",[39,1807,1808],{},"Is the service working right now?",[15,1810,1811],{},"Simple as that sounds, the implications are significant. Without a statuspage, users experiencing an outage have no choice but to contact support — via email, phone, or chat. The result: overwhelmed support teams, frustrated customers, and a loss of control over the narrative. With a well-maintained statuspage, incidents are communicated proactively before the first support ticket is filed.",[15,1813,1814],{},"Companies like GitHub, Cloudflare, and Stripe have operated public statuspages for years. But the concept is far from limited to large enterprises. Any organization that runs digital services — from early-stage SaaS startups to internal IT departments — benefits from a transparent statuspage.",[10,1816,1818],{"id":1817},"who-needs-a-statuspage","Who Needs a Statuspage?",[25,1820,1822],{"id":1821},"saas-companies-and-software-providers","SaaS Companies and Software Providers",[15,1824,1825],{},"For SaaS companies, a statuspage is not optional. Customers pay for availability. When a service goes down and the provider stays silent, uncertainty takes over. A statuspage communicates: We are aware, we are working on it, and we will keep you updated. That is the difference between a professional vendor and one that cannot be trusted.",[25,1827,1829],{"id":1828},"hosting-providers-and-infrastructure-companies","Hosting Providers and Infrastructure Companies",[15,1831,1832],{},"Hosting providers operate at a layer where outages cascade. When a data center has issues, hundreds of customers are potentially affected. A statuspage with granular breakdowns by region, service, and system is essential at this scale.",[25,1834,1836],{"id":1835},"e-commerce-and-online-retail","E-Commerce and Online Retail",[15,1838,1839],{},"In e-commerce, downtime means direct revenue loss. A statuspage informs not just end customers but also payment processors, logistics partners, and internal teams. During peak seasons — Black Friday, holiday shopping — proactive status communication becomes business-critical.",[25,1841,1843],{"id":1842},"internal-it-teams","Internal IT Teams",[15,1845,1846],{},"Not every statuspage needs to be public. Internal statuspages inform employees about the status of corporate systems: ERP, CRM, email, VPN, internal tools. This reduces internal support tickets and gives the IT department room to focus on resolution rather than communication.",[25,1848,1850],{"id":1849},"regulated-industries","Regulated Industries",[15,1852,1853],{},"Organizations in regulated industries — financial services, healthcare, government — often face compliance requirements around availability documentation and incident communication. A statuspage with a complete history and SLA tracking fulfills these requirements in a structured, auditable way.",[10,1855,1857],{"id":1856},"core-features-of-a-good-statuspage","Core Features of a Good Statuspage",[15,1859,1860],{},"Not all statuspages are created equal. The difference between a useful statuspage and a useless one comes down to features and execution.",[25,1862,1864],{"id":1863},"services-and-components","Services and Components",[15,1866,1867],{},"A statuspage should not display overall status as a single value. It should break down into logical services and components. A user relying on the API does not care about the dashboard status — and vice versa. Granularity creates relevance.",[25,1869,1871],{"id":1870},"real-time-status-display","Real-Time Status Display",[15,1873,1874],{},"The current status of each service must be visible at a glance. Common levels include: Operational, Degraded Performance, Partial Outage, and Major Outage. Color coding — green, yellow, orange, red — makes status intuitively readable without requiring users to parse text.",[25,1876,1878],{"id":1877},"incident-updates-with-timeline","Incident Updates with Timeline",[15,1880,1881],{},"During an incident, posting \"There is a problem\" once is not enough. Professional incident communication includes a complete timeline: detection, investigation, root cause identification, fix implementation, and resolution confirmation. Each step is documented with a timestamp.",[25,1883,1885],{"id":1884},"subscribers-and-notifications","Subscribers and Notifications",[15,1887,1888],{},"Not every user actively checks the statuspage. Subscriber functionality allows users to receive email notifications about incidents and maintenance. This is the difference between pull and push communication — and push wins in practice every time.",[25,1890,1892],{"id":1891},"scheduled-maintenance-windows","Scheduled Maintenance Windows",[15,1894,1895],{},"Maintenance is inevitable. A good statuspage allows teams to announce maintenance windows in advance, mark affected services, and notify subscribers ahead of time. This prevents unnecessary support tickets and demonstrates professional planning.",[25,1897,1899],{"id":1898},"uptime-history-and-sla-tracking","Uptime History and SLA Tracking",[15,1901,1902],{},"A statuspage without history is just a snapshot. Uptime history — ideally displayed as a calendar or timeline over 30, 60, or 90 days — gives users and prospective customers an overview of long-term reliability. For B2B buyers evaluating vendors, this is often a deciding factor.",[10,1904,1906],{"id":1905},"what-separates-a-good-statuspage-from-a-bad-one","What Separates a Good Statuspage from a Bad One?",[25,1908,1910],{"id":1909},"transparency-over-spin","Transparency Over Spin",[15,1912,1913],{},"The most common failure: statuspages that permanently display \"All Systems Operational\" while users are actively experiencing problems. A statuspage that hides or delays incident reports is worse than no statuspage at all. It actively erodes trust.",[15,1915,1916],{},"A good statuspage names problems clearly and promptly. It describes what is affected, what the impact is, and when a resolution is expected. Perfection is not the goal — honesty is.",[25,1918,1920],{"id":1919},"timeliness-and-speed","Timeliness and Speed",[15,1922,1923],{},"A statuspage that gets updated 30 minutes after an outage begins has missed its purpose. Updates should happen within minutes — ideally triggered automatically by the monitoring system that detected the issue.",[25,1925,1927],{"id":1926},"design-and-readability","Design and Readability",[15,1929,1930],{},"A statuspage must work at first glance. The user arrives with a specific question: Is my service running? The answer must be immediately visible — no scrolling, no searching, no interpretation required. Clean typography, consistent color coding, and a logical structure are non-negotiable.",[25,1932,1934],{"id":1933},"mobile-experience","Mobile Experience",[15,1936,1937],{},"Outages do not only happen during office hours. Users check status on their phones, on the go, in transit. A statuspage that performs poorly on mobile devices is practically unusable in real-world scenarios. Progressive Web App (PWA) support takes the mobile experience a step further, enabling home screen installation and offline access to the last known status.",[10,1939,1941],{"id":1940},"why-transparency-builds-trust","Why Transparency Builds Trust",[15,1943,1944],{},"Many companies hesitate to launch a public statuspage. The fear: if we make our outages public, we will lose customers. Reality consistently shows the opposite.",[25,1946,1948],{"id":1947},"outages-are-normal-silence-is-not","Outages Are Normal — Silence Is Not",[15,1950,1951],{},"Every service goes down eventually. Your customers know this. What customers will not accept is opacity. When a service is not working and the company stays silent, the trust damage extends far beyond the outage itself. Customers wonder: Do they even know there is a problem? Do they care?",[25,1953,1955],{"id":1954},"proactive-communication-reduces-support-load","Proactive Communication Reduces Support Load",[15,1957,1958],{},"Industry experience shows consistently that an actively maintained statuspage reduces support ticket volume during incidents by 30 to 50 percent. Instead of answering hundreds of identical tickets, a single statuspage entry directs all affected users to the same information.",[25,1960,1962],{"id":1961},"transparency-as-a-competitive-advantage","Transparency as a Competitive Advantage",[15,1964,1965],{},"In markets with many comparable providers, transparency becomes a differentiator. A vendor that openly displays uptime history signals confidence and professionalism. This matters especially for B2B buyers who evaluate vendors and want to assess reliability objectively.",[25,1967,1969],{"id":1968},"post-mortems-build-long-term-trust","Post-Mortems Build Long-Term Trust",[15,1971,1972],{},"The best organizations go a step further: after a major incident, they publish a detailed post-mortem. What happened? Why? What is being done to prevent recurrence? This level of openness builds more trust than a spotless uptime dashboard ever could.",[10,1974,1976],{"id":1975},"statuspage-vs-monitoring-why-you-need-both","Statuspage vs. Monitoring — Why You Need Both",[15,1978,1979],{},"A common mistake: monitoring and statuspages are treated as separate problems and solved with separate tools. In practice, this leads to friction, manual processes, and delayed updates.",[25,1981,1983],{"id":1982},"what-monitoring-does","What Monitoring Does",[15,1985,1986],{},"Monitoring watches the availability and performance of services. HTTP checks verify that a website is reachable. TCP checks test network services. Heartbeat checks monitor cron jobs and background processes. When a check fails, an alert is triggered.",[25,1988,1990],{"id":1989},"what-a-statuspage-does","What a Statuspage Does",[15,1992,1993],{},"A statuspage communicates status externally. It is the interface between the internal knowledge of an incident and the external communication to customers and stakeholders.",[25,1995,1997],{"id":1996},"the-gap-between-the-two","The Gap Between the Two",[15,1999,2000],{},"When monitoring and statuspage are separate systems, a gap emerges. Monitoring detects the outage, but the statuspage only gets updated manually — whenever someone remembers to do it. In the meantime — minutes or, in worst cases, hours — the statuspage falsely shows \"All Operational.\"",[15,2002,2003],{},"The solution is an integrated platform that combines monitoring and statuspage in one system. When a check fails, an incident is automatically created and the statuspage is updated. When the check recovers, the incident is automatically marked as resolved. No manual step, no delay.",[10,2005,2007],{"id":2006},"how-to-set-up-a-statuspage","How to Set Up a Statuspage",[25,2009,2011],{"id":2010},"step-1-define-your-services","Step 1: Define Your Services",[15,2013,2014],{},"Before the statuspage goes live, decide which services to display. A good structure follows the user perspective, not the internal architecture. Instead of \"PostgreSQL Primary\" and \"Redis Cluster,\" use: \"API,\" \"Dashboard,\" \"Email Delivery,\" \"Payments.\"",[25,2016,2018],{"id":2017},"step-2-set-up-monitoring","Step 2: Set Up Monitoring",[15,2020,2021],{},"Each service should have at least one monitoring check. HTTP checks for web applications, TCP checks for database connections, heartbeat checks for background processes. The checks must be representative — they should test exactly what the end user perceives as \"working\" or \"not working.\"",[25,2023,2025],{"id":2024},"step-3-define-your-incident-workflow","Step 3: Define Your Incident Workflow",[15,2027,2028],{},"Who can create incidents? Who writes updates? At what intervals are updates posted? These processes should be defined before the first outage occurs. A clear workflow with defined roles and escalation levels prevents chaos when it matters most.",[25,2030,2032],{"id":2031},"step-4-design-and-publish-the-statuspage","Step 4: Design and Publish the Statuspage",[15,2034,2035],{},"The design and branding of the statuspage should match the company. The statuspage is a customer touchpoint — it should look professional and reflect the brand identity. A subdomain (status.example.com) or custom domain are common approaches.",[25,2037,2039],{"id":2038},"step-5-enable-subscribers-and-spread-the-word","Step 5: Enable Subscribers and Spread the Word",[15,2041,2042],{},"After launch, customers need to know the statuspage exists. Links in the main website footer, documentation, onboarding emails, and support portal are standard placements. Subscriber options — email newsletter, RSS, webhooks — enable push notifications for those who want them.",[25,2044,2046],{"id":2045},"example-setting-up-a-statuspage-with-livck","Example: Setting Up a Statuspage with LIVCK",[15,2048,2049],{},"LIVCK is a monitoring and statuspage solution from Germany that combines both functions in a single platform. The setup process illustrates how an integrated approach works in practice:",[15,2051,2052,2055],{},[39,2053,2054],{},"Installation:"," LIVCK can be self-hosted via Docker Compose in minutes — on any server with Docker support. Alternatively, a managed service is available. A cloud option with a free starter plan is in the works.",[15,2057,2058,2061],{},[39,2059,2060],{},"Configure Monitoring:"," After installation, set up checks — HTTP(S) for web applications, TCP for network services, SSL for certificates, Heartbeat for cron jobs and scheduled tasks. Against false alarms, LIVCK checks from multiple independent locations — an incident is only triggered once the majority of probes confirm the outage.",[15,2063,2064,2067],{},[39,2065,2066],{},"Design the Statuspage:"," The drag-and-drop designer lets you build the statuspage visually. Three themes serve as starting points. Custom branding — logo, colors, domain — is included in every plan at no extra cost.",[15,2069,2070,2073],{},[39,2071,2072],{},"Incident Management:"," LIVCK uses a full incident workflow with Outage Linking. This means: when multiple services are affected by the same root cause, they are grouped into a single incident. Subscribers are automatically notified via email, Slack, Discord, Telegram, or SMS.",[15,2075,2076,2079],{},[39,2077,2078],{},"GDPR by Design:"," Since LIVCK can be self-hosted or runs in German data centers, all data stays under your control. For organizations with strict data privacy requirements or those operating in regulated industries, this is a decisive advantage over US-based providers.",[10,2081,2083],{"id":2082},"best-practices-for-ongoing-operations","Best Practices for Ongoing Operations",[15,2085,2086],{},"Setting up a statuspage is the first step. Operating it well over time is the real challenge.",[15,2088,2089,2092],{},[39,2090,2091],{},"Automate Where Possible:"," Manual statuspage updates are error-prone and slow. Every incident detected by monitoring should be automatically reflected on the statuspage without human intervention for the initial status change.",[15,2094,2095,2098],{},[39,2096,2097],{},"Communicate Maintenance Regularly:"," Even when there are no incidents, regular maintenance announcements show that the statuspage is actively maintained. A statuspage with no entries for months looks abandoned — and abandoned tools do not inspire confidence.",[15,2100,2101,2104],{},[39,2102,2103],{},"Use Clear Language:"," Incident updates should be written for end users, not for engineers. Instead of \"OOM kill on pod-xyz-123 in k8s cluster,\" write: \"Our dashboard is currently unavailable. We have identified the cause and are working on a fix. We expect to have this resolved within the next 30 minutes.\"",[15,2106,2107,2110],{},[39,2108,2109],{},"Grow Your Subscriber Base:"," The more users subscribed to the statuspage, the more effective incident communication becomes. Active prompts to subscribe during onboarding and in support interactions pay dividends over time.",[15,2112,2113,2116],{},[39,2114,2115],{},"Review and Improve:"," After every major incident, review how the statuspage communication went. Was the first update timely? Were updates frequent enough? Was the language clear? Continuous improvement of the incident communication process is just as important as improving the infrastructure itself.",[10,2118,564],{"id":563},[15,2120,2121],{},"A statuspage is not a nice-to-have — it is a core component of professional service operations. It reduces support load, builds trust through transparency, and gives customers confidence that incidents are detected and handled.",[15,2123,2124],{},"The greatest impact comes when the statuspage is integrated with the monitoring system. Separate tools for monitoring and statuspage create delays and manual overhead. An integrated approach — like LIVCK, which combines monitoring, incident management, and statuspage in a single platform — closes that gap entirely.",[15,2126,2127],{},"Whether public-facing for customers or internal for teams: anyone operating digital services needs a statuspage. Not as a marketing tool, but as an instrument for honest, timely communication. Because trust is not built through perfect uptime — it is built through the professional handling of the moments when things go wrong.",{"title":572,"searchDepth":573,"depth":573,"links":2129},[2130,2131,2138,2146,2152,2158,2163,2171,2172],{"id":613,"depth":573,"text":1799},{"id":1817,"depth":573,"text":1818,"children":2132},[2133,2134,2135,2136,2137],{"id":1821,"depth":578,"text":1822},{"id":1828,"depth":578,"text":1829},{"id":1835,"depth":578,"text":1836},{"id":1842,"depth":578,"text":1843},{"id":1849,"depth":578,"text":1850},{"id":1856,"depth":573,"text":1857,"children":2139},[2140,2141,2142,2143,2144,2145],{"id":1863,"depth":578,"text":1864},{"id":1870,"depth":578,"text":1871},{"id":1877,"depth":578,"text":1878},{"id":1884,"depth":578,"text":1885},{"id":1891,"depth":578,"text":1892},{"id":1898,"depth":578,"text":1899},{"id":1905,"depth":573,"text":1906,"children":2147},[2148,2149,2150,2151],{"id":1909,"depth":578,"text":1910},{"id":1919,"depth":578,"text":1920},{"id":1926,"depth":578,"text":1927},{"id":1933,"depth":578,"text":1934},{"id":1940,"depth":573,"text":1941,"children":2153},[2154,2155,2156,2157],{"id":1947,"depth":578,"text":1948},{"id":1954,"depth":578,"text":1955},{"id":1961,"depth":578,"text":1962},{"id":1968,"depth":578,"text":1969},{"id":1975,"depth":573,"text":1976,"children":2159},[2160,2161,2162],{"id":1982,"depth":578,"text":1983},{"id":1989,"depth":578,"text":1990},{"id":1996,"depth":578,"text":1997},{"id":2006,"depth":573,"text":2007,"children":2164},[2165,2166,2167,2168,2169,2170],{"id":2010,"depth":578,"text":2011},{"id":2017,"depth":578,"text":2018},{"id":2024,"depth":578,"text":2025},{"id":2031,"depth":578,"text":2032},{"id":2038,"depth":578,"text":2039},{"id":2045,"depth":578,"text":2046},{"id":2082,"depth":573,"text":2083},{"id":563,"depth":573,"text":564},"Fundamentals","What is a statuspage, why do you need one, and what makes a good statuspage? Everything about features, use cases, and why transparency builds trust.",[2176,2177,2179,2181],{"question":1799,"answer":1802},{"question":1818,"answer":2178},"Any organization running digital services benefits: SaaS companies, hosting and infrastructure providers, e-commerce shops, internal IT teams, and regulated industries such as financial services, healthcare, and government. Internal statuspages also inform employees about the status of corporate systems like ERP, CRM, email, and VPN.",{"question":1906,"answer":2180},"A good statuspage names problems clearly and promptly instead of permanently showing All Systems Operational, updates within minutes — ideally triggered automatically by the monitoring system — and is readable at first glance and on mobile. A statuspage that hides or delays incident reports is worse than none, because it actively erodes trust.",{"question":2007,"answer":2182},"In five steps: define your services from the user perspective, set up at least one monitoring check per service (HTTP, TCP, heartbeat), define an incident workflow with roles and escalation levels, design and publish the statuspage, and enable subscribers while spreading the word.",{},"\u002Fguides\u002Fen\u002Fwhat-is-a-statuspage","was-ist-eine-statuspage",10,[609,1785,1150],[611,1153],{"title":1794,"description":2174},"guides\u002Fen\u002Fwhat-is-a-statuspage","xziKDCNdJklhgCAHlV3BNgSAswcH6qkFGEX9AxCX3-8",1783478157477]