Real-Time Vision AI for Workplace Safety Compliance
A production-ready PPE detection service designed to cut manual safety monitoring by 70-80%

Executive Summary
Projected reduction in manual monitoring effort
Projected safety review time per shift, down from 2-4 hours
Image, recorded video, and real-time live stream analysis
Flagged events backed by annotated outputs and structured detection results
The Challenge
Industrial operators in manufacturing, construction, utilities, and warehousing have a continuous obligation to monitor workplace safety and PPE compliance across every site and shift. In practice, safety teams cannot watch every worker and every camera feed at all times. Most compliance programs still rely on floor walkthroughs, spot checks, CCTV review, manual logs, spreadsheets, and periodic audits. PPE violations such as missing helmets, missing vests, or unsafe worker behavior may happen briefly and disappear before a safety officer sees them. A typical site may have 8-16 CCTV feeds, but a safety officer can actively monitor only a few at a time, so compliance coverage is limited to spot checks and floor walkthroughs — often consuming 2-4 hours per shift — with audit evidence depending entirely on what humans happened to notice and log by hand.
Key Pain Points
- ✕Safety officers can only watch a few of 8-16 CCTV feeds at any time
- ✕Brief PPE violations disappear before anyone sees them
- ✕2-4 hours per shift spent on walkthroughs and manual CCTV review
- ✕Audit evidence limited to what humans happened to notice and log by hand
Our Solution
Amasa built a real-time vision AI service for workplace safety and PPE compliance monitoring — a capability case study with modeled impact. The core service is a FastAPI backend running RF-DETR, a real-time detection transformer model, supporting three input modes: single uploaded image detection, recorded video processed frame by frame with batching, and real-time live stream analysis. OpenCV and the supervision library handle frame processing and annotation, boto3 connects to S3 storage, and ffmpeg supports video transcoding. Inference returns bounding boxes and confidence scores; annotated outputs and structured detection results are exposed via REST API so alerts can be routed into dashboards or existing safety workflows, where officers review flagged events and act. The service ships production-ready with GPU Docker deployment, model warmup, structured logging, and monitoring, plus commercial-readiness work including a third-party license audit and GPL risk analysis. The architecture is on-prem-ready and extensible beyond PPE to intrusion detection, vehicle zones, and quality inspection.
Implementation Approach
Detection Scope & Model Selection
Defined PPE and safety detection targets and selected RF-DETR, a real-time detection transformer
Inference Service Build
Built a FastAPI service handling image, frame-by-frame video, and live stream inference with OpenCV annotation
Production Hardening
Added GPU Docker deployment, model warmup, structured logging, monitoring, and S3 integration
Commercial Readiness
Completed a third-party license audit and GPL risk analysis for safe commercial deployment
The Results
The implementation delivered transformative results across all key metrics, with immediate impact on operational efficiency, accuracy, and customer satisfaction.
Impact Metrics
| Metric | Before | After | Improvement |
|---|---|---|---|
| Monitoring model | Manual CCTV watching and spot checks | Automated detection across images, video, streams | Continuous coverage |
| Monitoring effort | 2-4 hours per shift | 15-30 minutes reviewing alerts (projected) | 70-80% reduction |
| Violation detection | Often found after the fact | Real-time detection on live streams | Immediate |
| Audit evidence | Manual logs and periodic records | Annotated outputs and structured results | Continuous audit trail |
Key Takeaways
- Real-time detection shifts safety teams from watching feeds to acting on alerts
- Annotated outputs and structured results turn compliance monitoring into a continuous audit trail
- GPU Docker packaging and on-prem readiness matter for industrial sites with data and network constraints
- License auditing and GPL risk analysis are essential groundwork before deploying open-source vision models commercially
Inside the Capability
The Workflow Before
- Safety officers walked the floor during shifts.
- Teams manually watched a small number of CCTV feeds.
- Violations were logged by hand in spreadsheets or paper records.
- Supervisors followed up after the fact.
- Audit evidence depended on what humans happened to notice.
A typical site may have 8–16 CCTV feeds, but a safety officer can actively monitor only a few at a time. Compliance coverage was limited to spot checks and floor walkthroughs – often 2–4 hours per shift – while brief PPE violations disappeared before anyone saw them.
The Workflow After
- A camera feed, uploaded video, or image is sent to the REST API.
- RF-DETR inference returns bounding boxes and confidence scores.
- Uploaded videos are processed frame by frame with batching.
- Live streams are analyzed in real time.
- Annotated outputs and structured detection results are made available through the API.
- Alerts are routed into dashboards or existing safety workflows.
- Safety officers review flagged events and take action.
Modeled Impact
This is a capability case study: the impact below is modeled for a typical multi-site industrial operator, not measured from a specific deployment.
| Metric | Before | After |
|---|---|---|
| Monitoring model | Manual CCTV watching and spot checks | Automated detection across images, video, and streams |
| Monitoring effort | 2–4 hours per shift | 15–30 minutes per shift reviewing alerts |
| Manual monitoring reduction | Baseline | 70–80% projected reduction |
| Violation detection | Often found after the fact | Real-time detection on live streams |
| Feed coverage | A few feeds watched intermittently | Connected site feeds monitored continuously |
| Audit evidence | Manual logs and periodic records | Annotated outputs and structured detection results |
Built for Production, Not a Demo
The service ships with GPU Docker deployment, model warmup, structured logging, and monitoring. Amasa also completed commercial-readiness work including a third-party license audit and GPL risk analysis – essential groundwork before deploying open-source vision models commercially. The architecture is on-prem-ready and extensible beyond PPE detection to intrusion detection, vehicle zones, and quality inspection.
The Takeaway
Amasa built a production-ready vision AI service that turns CCTV, images, and video streams into real-time compliance intelligence. For multi-site operators, the system is designed to reduce manual monitoring effort by 70–80%, shift safety teams from watching feeds to acting on alerts, and create a continuous audit trail.
Quick Facts
Industry
Manufacturing
Solution Type
Computer Vision
Published
August 24, 2026
Technologies Used
Related Resources
Ready to Achieve Similar Results?
Let's discuss how we can transform your business with AI solutions tailored to your specific needs
