Skip to content

Alert Rules Configuration

Overview

Advanced guide to configuring and managing alert rules in the SR Operations KPI module.

Last Updated: February 2026 Duration: 30 minutes Audience: Managers, System Administrators


Lesson 1: Alert System Architecture

Learning Objectives

  • Understand how alerts are generated
  • Configure effective alert rules
  • Avoid alert fatigue
  • Troubleshoot alert issues

How the Alert System Works

WIP Snapshot (every 15 min)
Rule Engine evaluates all active rules
For each rule that triggers:
    - Check cooldown period
    - If cooldown passed → Create Alert
    - Notify configured managers
    - Send SMS if enabled

Components

Component Purpose
WIP Snapshot Point-in-time data capture
Alert Rule Defines trigger conditions
Alert The actual notification created
Cooldown Prevents duplicate alerts

Lesson 2: Understanding Metrics

Available Metrics

These metrics can be monitored by alert rules:

WIP Metrics:

Metric ID Description Unit
da_wip_days Days of DA work in queue Days
laptop_wip_days Days of laptop work queued Days
hd_wip_days Days of HD work queued Days
laptop_queue Laptops waiting to process Count
hd_queue_drives Drives waiting to process Count

Schedule Metrics:

Metric ID Description Unit
schedule_fill_pct Next week schedule fullness Percentage
next_week_solid_days Days with 15+ stops Count (0-5)
today_stops Pickup stops today Count

Production Metrics:

Metric ID Description Unit
devices_processed_today Devices processed today Count

Metric Calculations

WIP Days:

WIP Days = Current Queue ÷ Daily Capacity

Schedule Fill:

Fill % = (Solid Days ÷ 5) × 100


Lesson 3: Condition Types

Comparison Operators

Operator Code Example
Greater Than gt Value > 6
Greater Than or Equal gte Value >= 6
Less Than lt Value < 3
Less Than or Equal lte Value <= 3
Equal To eq Value = 0

Choosing the Right Condition

Use Greater Than (>) for: - Backlog alerts (WIP too high) - Queue size warnings - Overtime indicators

Use Less Than (<) for: - Schedule fill warnings - Production below expectations - Understaffing indicators


Lesson 4: Severity Levels

When to Use Each Level

Severity Use When Expected Response
Critical Immediate action required Drop everything, address now
Warning Attention needed soon Address within hours
Info Awareness only Review when convenient

Severity Guidelines

Critical: - Customer commitments at risk - Safety concerns - Revenue impact > $1000/day - Multiple days of backlog

Warning: - Trending toward problems - Single day of excess backlog - Schedule filling slowly - Production slightly behind

Info: - Milestone reached - Unusual but not problematic - FYI situations


Lesson 5: Crafting Effective Rules

Rule Design Principles

1. Be Specific - Don't alert on everything - Focus on actionable situations - Include clear suggested actions

2. Set Appropriate Thresholds - Too low = alert fatigue - Too high = miss problems - Start conservative, adjust based on experience

3. Use Meaningful Messages - Include actual values: {value} - Include thresholds: {threshold} - Be concise but clear

4. Configure Cooldowns - Prevents repeated alerts for same issue - Match to realistic response time - Longer for persistent issues

Example Rules

Good Rule:

Name: DA Backlog Critical
Metric: da_wip_days
Condition: Greater Than
Threshold: 6
Severity: Critical
Message: "DA backlog at {value:.1f} days - exceeds {threshold} day limit"
Action: "Keep all DA techs focused on DA. Consider overtime."
Cooldown: 2 hours

Poor Rule:

Name: Any WIP
Metric: da_wip_days
Condition: Greater Than
Threshold: 0
Severity: Critical
Message: "There is WIP"
Action: (none)
Cooldown: 0

Problems: Threshold too low, no useful message, no action, no cooldown.


Lesson 6: Default Rules Explained

DA Backlog Rules

DA Backlog Critical (da_wip_days > 6) - Triggers when 6+ days of DA work waiting - This means major delays for customers - Action: All-hands-on-deck for DA

DA Backlog Warning (da_wip_days > 4) - Early warning before critical - Time to adjust staffing - Action: Prioritize DA work

Route Rules

Routes Critical (next_week_solid_days < 3) - Less than 3 solid days scheduled - Trucks will be underutilized - Action: Push call center to book

Schedule Fill Warning (schedule_fill_pct < 60%) - Schedule less than 60% full - Similar to routes critical but percentage-based - Action: Focus on booking

Queue Rules

HD Queue Critical (hd_queue_drives > 250) - 250+ drives waiting - At 100/day capacity, 2.5+ days behind - Action: Prioritize HD bins

Laptop Queue Warning (laptop_queue > 100) - 100+ laptops waiting - At 20/day, 5+ days of work - Action: Assign additional tech


Lesson 7: Managing Alert Fatigue

Signs of Alert Fatigue

  • Alerts being ignored
  • Critical alerts not addressed quickly
  • Staff complaining about notifications
  • High volume of acknowledged-without-action

Prevention Strategies

1. Right-Size Thresholds - Review triggered alerts weekly - If > 50% are false alarms, adjust thresholds - If problems aren't caught, lower thresholds

2. Use Appropriate Severity - Reserve Critical for true emergencies - Most alerts should be Warning - Use Info sparingly

3. Consolidate Similar Rules - Don't have 5 rules for the same issue - One rule per problem area is usually enough

4. Effective Cooldowns - Minimum 1 hour for warnings - Minimum 2 hours for critical - Longer for persistent issues (4-8 hours)

Alert Review Process

Weekly: 1. Review all alerts from past week 2. Check which were actionable 3. Note any missed problems 4. Adjust rules as needed

Monthly: 1. Analyze alert patterns 2. Remove ineffective rules 3. Add rules for new problem areas 4. Update thresholds based on data


Lesson 8: Troubleshooting

Common Issues

"Alert didn't fire when it should have" 1. Check rule is active 2. Verify metric is being captured in snapshots 3. Check cooldown hasn't blocked it 4. Verify threshold is correct

"Getting too many alerts" 1. Increase threshold 2. Increase cooldown 3. Change severity to Info 4. Consider disabling rule

"Alert message shows wrong values" 1. Check {value} and {threshold} placeholders 2. Verify format strings (e.g., {value:.1f} for decimals)

"SMS not sending" 1. Check employee has mobile_phone_sms set 2. Verify Twilio configuration (sr_twilio_sms module) 3. Check SMS log for errors

Viewing Alert History

SR Operations → WIP and Alerts → All Alerts

Filter by: - Rule (which rule triggered) - Date range - Acknowledged status - Severity

Testing Rules

To test a rule without waiting: 1. Create a manual WIP snapshot with test values 2. Manually run the rule engine (Technical → Scheduled Actions) 3. Check if alert was created 4. Delete test data when done


Quick Reference

Rule Creation Checklist

  • Clear, descriptive name
  • Appropriate metric selected
  • Threshold tested against real data
  • Severity matches urgency
  • Message includes {value} and {threshold}
  • Suggested action is specific
  • Cooldown is reasonable (1-8 hours)
  • Managers assigned if needed
  • SMS enabled if appropriate
Situation Cooldown
True emergency 1-2 hours
Daily concern 4 hours
Persistent issue 8 hours
Weekly review 24 hours

Metric Quick Reference

Short Name Full Name Typical Range
da_wip DA WIP Days 0-10
laptop_wip Laptop WIP Days 0-7
hd_wip HD WIP Days 0-5
fill Schedule Fill % 0-100
solid Solid Days 0-5

Questions? Contact IT support or refer to the technical documentation.