Alert Rules Configuration¶
Overview¶
Advanced guide to configuring and managing alert rules in the SR Operations KPI module.
Last Updated: February 2026 Duration: 30 minutes Audience: Managers, System Administrators
Lesson 1: Alert System Architecture¶
Learning Objectives¶
- Understand how alerts are generated
- Configure effective alert rules
- Avoid alert fatigue
- Troubleshoot alert issues
How the Alert System Works¶
WIP Snapshot (every 15 min)
↓
Rule Engine evaluates all active rules
↓
For each rule that triggers:
- Check cooldown period
- If cooldown passed → Create Alert
- Notify configured managers
- Send SMS if enabled
Components¶
| Component | Purpose |
|---|---|
| WIP Snapshot | Point-in-time data capture |
| Alert Rule | Defines trigger conditions |
| Alert | The actual notification created |
| Cooldown | Prevents duplicate alerts |
Lesson 2: Understanding Metrics¶
Available Metrics¶
These metrics can be monitored by alert rules:
WIP Metrics:
| Metric ID | Description | Unit |
|---|---|---|
da_wip_days |
Days of DA work in queue | Days |
laptop_wip_days |
Days of laptop work queued | Days |
hd_wip_days |
Days of HD work queued | Days |
laptop_queue |
Laptops waiting to process | Count |
hd_queue_drives |
Drives waiting to process | Count |
Schedule Metrics:
| Metric ID | Description | Unit |
|---|---|---|
schedule_fill_pct |
Next week schedule fullness | Percentage |
next_week_solid_days |
Days with 15+ stops | Count (0-5) |
today_stops |
Pickup stops today | Count |
Production Metrics:
| Metric ID | Description | Unit |
|---|---|---|
devices_processed_today |
Devices processed today | Count |
Metric Calculations¶
WIP Days:
Schedule Fill:
Lesson 3: Condition Types¶
Comparison Operators¶
| Operator | Code | Example |
|---|---|---|
| Greater Than | gt |
Value > 6 |
| Greater Than or Equal | gte |
Value >= 6 |
| Less Than | lt |
Value < 3 |
| Less Than or Equal | lte |
Value <= 3 |
| Equal To | eq |
Value = 0 |
Choosing the Right Condition¶
Use Greater Than (>) for: - Backlog alerts (WIP too high) - Queue size warnings - Overtime indicators
Use Less Than (<) for: - Schedule fill warnings - Production below expectations - Understaffing indicators
Lesson 4: Severity Levels¶
When to Use Each Level¶
| Severity | Use When | Expected Response |
|---|---|---|
| Critical | Immediate action required | Drop everything, address now |
| Warning | Attention needed soon | Address within hours |
| Info | Awareness only | Review when convenient |
Severity Guidelines¶
Critical: - Customer commitments at risk - Safety concerns - Revenue impact > $1000/day - Multiple days of backlog
Warning: - Trending toward problems - Single day of excess backlog - Schedule filling slowly - Production slightly behind
Info: - Milestone reached - Unusual but not problematic - FYI situations
Lesson 5: Crafting Effective Rules¶
Rule Design Principles¶
1. Be Specific - Don't alert on everything - Focus on actionable situations - Include clear suggested actions
2. Set Appropriate Thresholds - Too low = alert fatigue - Too high = miss problems - Start conservative, adjust based on experience
3. Use Meaningful Messages
- Include actual values: {value}
- Include thresholds: {threshold}
- Be concise but clear
4. Configure Cooldowns - Prevents repeated alerts for same issue - Match to realistic response time - Longer for persistent issues
Example Rules¶
Good Rule:
Name: DA Backlog Critical
Metric: da_wip_days
Condition: Greater Than
Threshold: 6
Severity: Critical
Message: "DA backlog at {value:.1f} days - exceeds {threshold} day limit"
Action: "Keep all DA techs focused on DA. Consider overtime."
Cooldown: 2 hours
Poor Rule:
Name: Any WIP
Metric: da_wip_days
Condition: Greater Than
Threshold: 0
Severity: Critical
Message: "There is WIP"
Action: (none)
Cooldown: 0
Problems: Threshold too low, no useful message, no action, no cooldown.
Lesson 6: Default Rules Explained¶
DA Backlog Rules¶
DA Backlog Critical (da_wip_days > 6) - Triggers when 6+ days of DA work waiting - This means major delays for customers - Action: All-hands-on-deck for DA
DA Backlog Warning (da_wip_days > 4) - Early warning before critical - Time to adjust staffing - Action: Prioritize DA work
Route Rules¶
Routes Critical (next_week_solid_days < 3) - Less than 3 solid days scheduled - Trucks will be underutilized - Action: Push call center to book
Schedule Fill Warning (schedule_fill_pct < 60%) - Schedule less than 60% full - Similar to routes critical but percentage-based - Action: Focus on booking
Queue Rules¶
HD Queue Critical (hd_queue_drives > 250) - 250+ drives waiting - At 100/day capacity, 2.5+ days behind - Action: Prioritize HD bins
Laptop Queue Warning (laptop_queue > 100) - 100+ laptops waiting - At 20/day, 5+ days of work - Action: Assign additional tech
Lesson 7: Managing Alert Fatigue¶
Signs of Alert Fatigue¶
- Alerts being ignored
- Critical alerts not addressed quickly
- Staff complaining about notifications
- High volume of acknowledged-without-action
Prevention Strategies¶
1. Right-Size Thresholds - Review triggered alerts weekly - If > 50% are false alarms, adjust thresholds - If problems aren't caught, lower thresholds
2. Use Appropriate Severity - Reserve Critical for true emergencies - Most alerts should be Warning - Use Info sparingly
3. Consolidate Similar Rules - Don't have 5 rules for the same issue - One rule per problem area is usually enough
4. Effective Cooldowns - Minimum 1 hour for warnings - Minimum 2 hours for critical - Longer for persistent issues (4-8 hours)
Alert Review Process¶
Weekly: 1. Review all alerts from past week 2. Check which were actionable 3. Note any missed problems 4. Adjust rules as needed
Monthly: 1. Analyze alert patterns 2. Remove ineffective rules 3. Add rules for new problem areas 4. Update thresholds based on data
Lesson 8: Troubleshooting¶
Common Issues¶
"Alert didn't fire when it should have" 1. Check rule is active 2. Verify metric is being captured in snapshots 3. Check cooldown hasn't blocked it 4. Verify threshold is correct
"Getting too many alerts" 1. Increase threshold 2. Increase cooldown 3. Change severity to Info 4. Consider disabling rule
"Alert message shows wrong values" 1. Check {value} and {threshold} placeholders 2. Verify format strings (e.g., {value:.1f} for decimals)
"SMS not sending" 1. Check employee has mobile_phone_sms set 2. Verify Twilio configuration (sr_twilio_sms module) 3. Check SMS log for errors
Viewing Alert History¶
Filter by: - Rule (which rule triggered) - Date range - Acknowledged status - Severity
Testing Rules¶
To test a rule without waiting: 1. Create a manual WIP snapshot with test values 2. Manually run the rule engine (Technical → Scheduled Actions) 3. Check if alert was created 4. Delete test data when done
Quick Reference¶
Rule Creation Checklist¶
- Clear, descriptive name
- Appropriate metric selected
- Threshold tested against real data
- Severity matches urgency
- Message includes {value} and {threshold}
- Suggested action is specific
- Cooldown is reasonable (1-8 hours)
- Managers assigned if needed
- SMS enabled if appropriate
Recommended Cooldowns¶
| Situation | Cooldown |
|---|---|
| True emergency | 1-2 hours |
| Daily concern | 4 hours |
| Persistent issue | 8 hours |
| Weekly review | 24 hours |
Metric Quick Reference¶
| Short Name | Full Name | Typical Range |
|---|---|---|
| da_wip | DA WIP Days | 0-10 |
| laptop_wip | Laptop WIP Days | 0-7 |
| hd_wip | HD WIP Days | 0-5 |
| fill | Schedule Fill % | 0-100 |
| solid | Solid Days | 0-5 |
Questions? Contact IT support or refer to the technical documentation.