A Vision-Language Model (VLM) powered aerial image analysis system for autonomous anomaly detection and site monitoring using drone imagery.
Aerial Site Intelligence ingests drone sensor imagery, runs structured scene understanding using GPT-4o Vision and AWS Rekognition, and surfaces anomaly alerts to operators. The system features:
- Vision-Language Model (VLM) Pipeline: GPT-4o Vision for scene understanding and semantic analysis
- Multi-Modal Detection: AWS Rekognition for object detection + SageMaker inference endpoints
- Persistent Agent Memory: Autonomous tracking of site-state evolution across missions
- Async Processing: Celery task queue for high-throughput image processing
- Anomaly Reasoning: Historical context-aware deviation detection
- RESTful API: Complete API for image management and analysis
┌─────────────────────────────────────────────────────────────┐
│ Drone Image Input │
└────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Django REST API & Upload │
└────────────────────────┬────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Celery Async Task Queue │
└────────────┬────────────────┬────────────────┬──────────────┘
│ │ │
┌────────▼────┐ ┌────────▼────┐ ┌───────▼────┐
│ Rekognition │ │ GPT-4 Vision│ │ SageMaker │
│ Detection │ │ Analysis │ │ Inference │
└────────┬────┘ └────────┬────┘ └───────┬────┘
│ │ │
└────────┬───────┴────────┬───────┘
│ │
▼ ▼
┌──────────────────┐ ┌──────────────┐
│ PostgreSQL DB │ │ Agent Memory │
└──────────────────┘ └──────────────┘
│
▼
┌──────────────────┐
│ Anomaly Detection│
└────────┬─────────┘
│
▼
┌──────────────────┐
│ Alert System │
└──────────────────┘
- Backend: Django 4.2, Django REST Framework
- Database: PostgreSQL
- Cache/Queue: Redis, Celery
- AI/ML: GPT-4o Vision, AWS Rekognition, AWS SageMaker
- Infrastructure: Docker, Docker Compose, AWS CloudFormation
- Deployment: Gunicorn, NGINX, AWS ECS/Lambda
- Python 3.11+
- PostgreSQL 15+
- Redis 7+
- Docker & Docker Compose
- AWS Account with:
- Rekognition access
- SageMaker endpoint
- S3 bucket
- Appropriate IAM roles
- OpenAI API key (for GPT-4o Vision)
- Clone the repository
git clone <repository-url>
cd ariealsiteintelligence- Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate- Install dependencies
pip install -r requirements.txt- Configure environment
cp .env.example .env
# Edit .env with your configuration- Run migrations
python manage.py migrate- Create superuser
python manage.py createsuperuser- Start development server
python manage.py runserver- Build and start services
docker-compose up -d- Run migrations
docker-compose exec django python manage.py migrate- Create superuser
docker-compose exec django python manage.py createsuperuserAccess the application at http://localhost:8000
Key environment variables (see .env.example):
# Django
DEBUG=False
SECRET_KEY=your-secret-key
ALLOWED_HOSTS=localhost,127.0.0.1
# Database
DATABASE_NAME=aerial_site_intelligence
DATABASE_USER=postgres
DATABASE_PASSWORD=postgres
DATABASE_HOST=localhost
DATABASE_PORT=5432
# AWS
AWS_ACCESS_KEY_ID=your-key
AWS_SECRET_ACCESS_KEY=your-secret
AWS_REGION=us-east-1
AWS_S3_BUCKET=aerial-site-intelligence
# OpenAI
OPENAI_API_KEY=your-api-key
GPT4_MODEL=gpt-4-vision-preview
# Celery
CELERY_BROKER_URL=redis://localhost:6379/0
CELERY_RESULT_BACKEND=redis://localhost:6379/0
# Application
ANOMALY_DETECTION_THRESHOLD=0.75http://localhost:8000/api/
GET /sites/- List all sitesPOST /sites/- Create new siteGET /sites/{id}/- Get site detailsGET /sites/{id}/stats/- Get site statisticsGET /sites/{id}/memory_context/- Get agent memory context
GET /missions/- List missionsPOST /missions/- Create missionGET /missions/{id}/- Get mission detailsPOST /missions/{id}/complete_mission/- Complete mission
GET /images/- List imagesPOST /images/- Upload imageGET /images/{id}/- Get image detailsPOST /images/{id}/reprocess/- Reprocess image
GET /detections/- List detectionsGET /detections/{id}/- Get detection details
GET /anomalies/- List anomaliesPOST /anomalies/{id}/acknowledge/- Acknowledge anomalyPOST /anomalies/{id}/resolve/- Resolve anomaly
GET /alerts/- List alertsGET /alerts/{id}/- Get alert details
GET /memory/- List memoriesGET /memory/{id}/- Get memory details
Upload an image:
curl -X POST http://localhost:8000/api/images/ \
-F "image=@image.jpg" \
-F "mission=1" \
-F "captured_at=2024-01-01T12:00:00Z"Get site statistics:
curl http://localhost:8000/api/sites/1/stats/Acknowledge anomaly:
curl -X POST http://localhost:8000/api/anomalies/1/acknowledge/-
Image Upload
- Image uploaded via API
- Stored in PostgreSQL and S3
- Queued for processing
-
Parallel Detection
- AWS Rekognition: Object/defect detection
- GPT-4o Vision: Scene understanding & semantic analysis
- SageMaker: Custom model inference
-
Analysis
- Store detections in database
- Generate semantic captions
- Extract scene context
-
Anomaly Detection
- Compare against agent memory
- Identify deviations from historical patterns
- Generate anomaly scores
-
Alerting
- Create alerts for anomalies
- Send notifications to operators
- Update agent memory with findings
The persistent agent memory tracks:
- Observations: Recent site observations and state
- Patterns: Identified trends and recurring anomalies
- Decisions: Historical decisions and actions taken
- Context: General site context and metadata
This enables the agent to:
- Reason over historical site evolution
- Detect subtle deviations from expected patterns
- Autonomously flag deviations without human review
- Provide context for operator decisions
Background tasks include:
process_aerial_image- Main image processing orchestrationrun_rekognition_detection- AWS Rekognition detectionrun_structural_analysis- Structural defect analysisrun_vlm_analysis- GPT-4o Vision analysisdetect_anomalies- Anomaly detection and reasoningsend_anomaly_alert- Alert notificationupdate_site_context- Update agent memorycleanup_expired_memories- Clean up expired memories
Monitor Celery tasks:
# In Docker
docker-compose exec celery celery -A aerial_site_intelligence inspect active
# View logs
docker-compose logs -f celery- Create CloudFormation stack
aws cloudformation create-stack \
--stack-name aerial-site-intelligence \
--template-body file://deploy/cloudformation-template.yaml- Deploy to ECS
# Build and push Docker image
docker build -t aerial-site-intelligence:latest .
docker tag aerial-site-intelligence:latest 123456789.dkr.ecr.us-east-1.amazonaws.com/aerial-site-intelligence:latest
docker push 123456789.dkr.ecr.us-east-1.amazonaws.com/aerial-site-intelligence:latest- Set up Lambda for S3 events
- Upload
deploy/lambda_handler.py - Configure S3 trigger
- Set environment variables
- Upload
- Set
DEBUG=False - Use strong
SECRET_KEY - Configure secure ALLOWED_HOSTS
- Set up HTTPS/SSL
- Configure database backups
- Set up monitoring and alerts
- Configure rate limiting
- Set up log aggregation
- Test disaster recovery
- Configure auto-scaling
# Run tests
python manage.py test
# With coverage
coverage run --source='.' manage.py test
coverage report- Image processing is fully asynchronous via Celery
- Redis caching for frequently accessed data
- Database query optimization with indexes
- S3 for scalable image storage
- SageMaker for efficient batch inference
- Django admin at
/admin/ - Celery flower (optional):
pip install flower && celery -A aerial_site_intelligence flower - CloudWatch logs for AWS resources
- PostgreSQL query logs
- Environment variables for sensitive data
- CORS configuration for API access
- Database encryption
- S3 bucket versioning and encryption
- IAM roles for AWS access
- API token authentication
# Check Redis connection
redis-cli ping
# Check Celery worker status
docker-compose ps
docker-compose logs celery
# Restart Celery
docker-compose restart celery# Check PostgreSQL
docker-compose exec db psql -U postgres -c "SELECT 1"
# Check migrations
python manage.py showmigrations- Verify AWS credentials in .env
- Check IAM permissions
- Verify service endpoints and regions
- Check CloudWatch logs
- Fork the repository
- Create a feature branch
- Make changes and add tests
- Submit a pull request
MIT License - see LICENSE file for details
For issues and questions:
- Create an GitHub issue
- Check documentation
- Contact development team
- Web dashboard for operators
- Mobile app for field teams
- Advanced pattern recognition
- Multi-site aggregation
- Custom ML model training
- Real-time video stream processing
- Integration with third-party services