Loading ...
Quick Summary

Practical, testable failover steps you can run before your next high-volume campaign — designed for Malaysian PBX admins using SIP trunks.

  • 7 clear actions (DNS, OPTIONS, SBC, routing, monitoring, number failover, test calls) that cut average outage impact from minutes to seconds when implemented properly.
  • Malaysia context: fixed‑broadband penetration was 64.1 per 100 premises (MCMC/DOSM dataset, June 2025), so ISP diversity and DNS design matter for local resilience.

Why this checklist matters: voice remains business‑critical in Malaysia even as data grows — outages translate to lost appointments, missed payments and frustrated customers. Use this checklist to make your SIP trunking resilient, measurable and auditable.

You’re the PBX admin who must keep agents live when campaigns run or when the office internet trips. SIP trunking gives you flexibility and lower costs, but it also introduces single points of failure if the design is opinionated around one ISP, one trunk registration, or a brittle DNS/TLS setup. This guide is focused on actionable configuration, testing, and operational controls you can finish in a maintenance window (or a coffee break) — with specific notes for Malaysian operators and examples that match common ITGTEL SIP Trunking plans.

What is the fastest, most reliable way to detect a SIP trunk outage?

Straight answer: instrument health checks at three layers — SIP signalling (OPTIONS/REGISTER), transport (ICMP/TCP), and application (synthetic test calls). Combine short OPTIONS polling (5–10s) with an SBC that interprets failure codes and an external monitor that alerts within 30 seconds. This multi‑layer detection avoids false positives and speeds failover decisions.

Technical detail: configure your PBX/SBC to send OPTIONS every 5–10 seconds to the carrier’s SIP proxy; reduce DNS TTLs for NAPTR/SRV/A records to 60–120s during high‑availability windows; terminate TLS at the SBC where possible so certificate issues don’t silently break registrations.

Further reading: 3CX — SIP Trunk Failover

A 7-step failover checklist PBX admins can run in 15 minutes

Direct answer: these seven items cover detection, routing, redundancy, and verification — follow them in order and mark each as PASS/FAIL during your test. Done right, they reduce call loss during trunk or ISP outages from minutes to seconds.

  1. Confirm multi‑path internet connectivity.

    Use at least two independent ISPs (different last‑mile providers) or an SD‑WAN with independent uplinks. Verify different AS paths with a traceroute to your SIP provider’s signalling IPs and confirm the two uplinks do not share local aggregation points.

  2. Set up multiple SIP registrations / trunk endpoints.

    Register your PBX/SBC to the primary carrier and a secondary carrier (or a secondary POP from the same carrier). If the carrier supports IP‑auth only, register the secondary via the carrier’s alternate IPs or use credentialed trunks as a backup.

  3. Tune health checks and failover thresholds.

    Configure OPTIONS heartbeat to 5–10s, with a 3‑strike failure policy before failing over (e.g., 3 x 5s = 15s). Map SIP response codes: 404/408/503 → immediate failover; transient 486/480 → retry. Keep DNS TTLs low (60–120s) during cutover windows.

  4. Implement an SBC (Session Border Controller) with clear routing rules.

    Use an SBC to centralise policy: prefer primary trunk, fallback to secondary, and finally route to a cloud PBX queue or voice blast service. The SBC should support translation, TLS/SRTP, SIP normalization, and per‑trunk call limits.

  5. Failover your DDI/DID mapping and emergency routing.

    Plan DID failover so inbound numbers can be rehomed to alternate queues or IVRs within your PBX routing logic. Update your emergency numbers routing policy if your secondary trunk’s geographic origination differs.

  6. Automate alerting and synthetic test calls.

    When OPTIONS fail, trigger an automated test call from an external monitor (cloud probe) to confirm audio path. Alert both network and contact‑centre ops via SMS/WhatsApp and escalate if the probe fails three times in a row.

  7. Run a live failover drill and document the result.

    Simulate the primary trunk outage during off‑peak hours: disable the primary registration, measure switchover time, record missed calls, and file a post‑mortem. Keep results in your runbook and repeat quarterly.

Quick check: if your measured switchover time is >60 seconds, re-evaluate DNS TTLs, OPTIONS interval, and whether your SBC is the decision-maker — these three elements typically dominate latency.

Which failover method is best for busy call centres using Malaysia SIP trunks?

Direct answer: busy call centres should prefer active‑active registrations (or registered SIP peers across multiple POPs) combined with an SBC that performs per‑call routing and dynamic load balancing. Active‑active removes the registration single‑point and lets carriers split traffic so failover is seamless.

For very high concurrency, combine active‑active trunks with call distribution logic on the PBX and persistent session handoff at the SBC. If active‑active isn’t available, use active‑hot standby with sub‑20s OPTIONS-based detection and pre‑warmed call queues on the backup.

Vendor reference on strategies: Cisco — High Availability and Network Design

How do DDI failover and number continuity work in practice?

Direct answer: DID continuity requires pre‑configured DID failover rules with your SIP provider and internal PBX mapping. The provider must support rehoming the DID (move the routing) to a secondary trunk or to a SIP URI you control; your PBX then maps that incoming URI to the correct queue or IVR.

Operational note: confirm with your Malaysia SIP trunk provider whether DDI rehoming is instant or needs a carrier update window. For high‑risk numbers (finance, bookings), pre‑configure voiced‑based fallback (IVR + recorded message) so calls do not drop during provider reconfiguration.

What monitoring and KPI targets should PBX teams set?

Direct answer: monitor three KPIs — registration health (OPTIONS pass rate), median failover time (seconds), and inbound call completion rate (%) during failovers. Target: OPTIONS pass rate 100% in normal state, median failover time ≤20s, and inbound call completion ≥95% during failover drills.

Add a synthetic call probe every 60s from an external cloud monitor located in Malaysia to validate audio path; log all failovers and keep a dashboard of last‑mile outages by ISP. Make SLA breach thresholds explicit for vendor escalation.

Common mistakes that make SIP failover brittle

Direct answer: the usual problems are single-ISP dependency, long DNS TTLs, relying only on REGISTER failures to detect outages, and no synthetic tests that verify RTP media. Avoid these and you eliminate most common outages.

  • Relying on one registration endpoint (single point of failure).
  • Missing SBC policies so SIP 503/408 responses are treated the same as permanent errors.
  • Assuming the WAN is fine because pings pass — voice quality needs RTP path checks.
Operator note

“In Malaysia the right strategy is ISP diversity + pre-warmed backup routes. Because fixed broadband penetration is not uniform, your PBX must assume at least one local link can fail.” — ITGTEL operations

How ITGTEL helps with SIP trunk failover in Malaysia

Direct answer: ITG Telecommunications Sdn Bhd (ITGTEL) provides SIP Trunking plans with multi‑channel capacity, DDI number packs and managed failover options — including pre‑built routing rules and SBC provisioning — so you can offload the carrier side of failover to an MCMC‑licensed provider with local operation and support.

If you run high‑volume voice operations, consider plans with multiple voice channels and pre‑reserved DDI blocks (for example, ITGTEL’s Universe and Space plans) so inbound numbers can be remapped quickly during an incident. See our SIP Trunking service page and the Universe plan for plan details and provisioning options.

Internal resources: ITGTEL SIP Trunking service page, SIP Trunking – Universe – 12 Month, and our operational guide SIP Trunking – Cutting Telecom Costs for Businesses in Malaysia.

Compliance/PDPA note: when you re-route DDI or play recorded IVR messages during failover, ensure any customer personal data captured is handled per PDPA — document the change of data flow and retention for audits.

Step‑by‑step quick checklist (one‑page copy for your runbook)

This is the one‑page snippet you can paste into an incident runbook and follow under pressure.

ActionCommand/CheckTarget
1. ISP diversitytraceroute to sip.pop1.example, sip.pop2.example2 independent AS paths
2. OPTIONS heartbeatOPTIONS interval 5–10s; 3 failures → failoverFailover ≤20s
3. SBC routingPrimary → Secondary → CloudPBXClear precedence and limits
4. DID rehomingVerify carrier support; pre‑staged IVRInbound continuity ≥95%
5. Synthetic test callExternal probe call every 60sAudio pass/fail logged
6. AlertingWhatsApp/SMS & email escalationOps alerted < 60s
7. DrillQuarterly simulation, record resultsPost‑mortem within 48h

Tip: include a pre‑approved message template (WhatsApp or SMS) for customer notification during long outages; this reduces support churn and preserves trust.

Data context: Malaysia fixed‑broadband penetration was 64.1 per 100 premises (Communication & Multimedia Dashboard via Department of Statistics Malaysia, as at June 2025). DOSM — TMP_MDE 2025 (source: MCMC Facts & Figures)

FAQ

How often should I run a failover drill?

Run a light drill (registration failover) monthly and a full simulated outage quarterly. Document switchover times, missed calls, and vendor response times for each drill.

Can I rely only on DNS SRV records for trunk failover?

DNS SRV helps but is slow if TTLs are high and it doesn’t detect media/RTP problems. Combine SRV with OPTIONS polling, multiple registrations and an SBC decision layer for robust failover.

Will ITGTEL waive registration fees if I move to a monthly SIP plan?

ITG Telecommunications Sdn Bhd waives the RM500 SIP registration fee on subscription for SIP Trunking plans. Contact our team for plan details and provisioning timelines.

Recommended engineering reading: 3CX – SIP Trunk Failover and Cisco – High Availability and Network Design.