Login Get started

Architecture

Understanding Hay's system architecture and design decisions

System Architecture#

Hay is designed as a modular, event-driven platform that scales with your needs. This document explains the key architectural decisions and how components work together.

High-Level Overview#

graph LR
  Clients["fa:fa-globe Clients<br/>Web / Mobile"]
  Gateway["fa:fa-shield-halved API Gateway<br/>Express"]
  Services["fa:fa-puzzle-piece Services<br/>Plugins"]
  Queue["fa:fa-list-check Message Queue<br/>RabbitMQ + Redis"]
  DB["fa:fa-database Database<br/>PostgreSQL"]

  Clients --> Gateway --> Services
  Gateway --> Queue
  Services --> DB

  style Clients fill:#e8f3ff,stroke:#568aff,color:#0a155c
  style Gateway fill:#e8f3ff,stroke:#568aff,color:#0a155c
  style Services fill:#e8f3ff,stroke:#568aff,color:#0a155c
  style Queue fill:#f5f5f5,stroke:#d4d4d4,color:#404040
  style DB fill:#f5f5f5,stroke:#d4d4d4,color:#404040

Core Components#

1. API Gateway#

The entry point for all client requests:

  • Authentication: JWT-based auth with refresh tokens
  • Rate Limiting: Prevents abuse and ensures fair usage
  • Request Validation: Schema validation using Zod
  • Response Formatting: Consistent API responses

2. Service Layer#

Business logic organized as modular services:

  • Conversation Service: Manages conversations and messages
  • Playbook Service: Handles playbook-based workflows and automation
  • Plugin Manager Service: Connects to external platforms via plugin instances
  • Analytics Service: Processes and stores metrics

3. Plugin System#

Hay's extensibility mechanism:

Plugins are defined using defineHayPlugin() from @hay/plugin-sdk. The plugin contract is HayPluginManifest in /server/types/plugin-sdk.types.ts, exposing capabilities via a runtime /metadata endpoint and interacting with the platform via HayGlobalContext.

Each plugin can:

  • Register hooks (onInitialize, onStart, onConnected, etc.)
  • Expose HTTP routes and MCP tools
  • Add new UI components
  • Access core services

4. Real-Time & Background Messaging#

There is no single event bus module. Real-time events are published to Redis pub/sub channels via ConversationEventsService. Background processing tasks are queued via RabbitMQ (rabbitmqService) and a Redis-backed JobQueueService.

5. Data Layer#

Persistent storage with caching:

  • PostgreSQL: Primary data store for conversations, users, settings
  • Redis: Caching layer and pub/sub for real-time features
  • File Storage: Attachments and media files (local filesystem by default, optional S3)

Data Flow#

Incoming Message#

  1. Message arrives via channel plugin (WhatsApp, email, webchat, etc.)
  2. Plugin persists the message and queues an orchestrator job via RabbitMQ
  3. Orchestrator worker picks up the job (perception → retrieval → execution layers)
  4. AI response is generated, guardrails are applied, and the reply is stored
  5. Real-time updates are published via ConversationEventsService (Redis pub/sub) to connected clients
  6. For channel conversations, ChannelDeliveryService automatically forwards the reply to the originating plugin's delivery endpoint

Scalability Considerations#

Horizontal Scaling#

Hay is designed to scale horizontally:

  • Stateless API servers: Scale by adding more instances
  • Background workers: Scale job processing independently
  • Database connection pooling: Efficient query distribution

Caching Strategy#

Multi-layer caching reduces database load:

graph LR
  A["fa:fa-browser Client Cache"] --> B["fa:fa-bolt Redis"] --> C["fa:fa-database Database"]

  style A fill:#e8f3ff,stroke:#568aff,color:#0a155c
  style B fill:#e8f3ff,stroke:#568aff,color:#0a155c
  style C fill:#f5f5f5,stroke:#d4d4d4,color:#404040
  • Client: Browser cache for static assets
  • Asset Domain: Optional configurable domain for serving static assets (not a CDN caching layer)
  • Redis: In-memory cache for hot data
  • Database: Source of truth

Message Queue#

RabbitMQ (via amqplib) handles orchestrator messaging, and a custom Postgres-backed job queue (using SKIP LOCKED) coordinated through Redis pub/sub handles background jobs:

  • Failed jobs are marked as FAILED by JobQueueService.failJob(); there is no automatic retry
  • Job priority is ordered via a SQL column, not a Bull-style priority queue
  • No rate limiting per job type

Security Architecture#

Defense in Depth#

Multiple security layers:

  1. Network: TLS/SSL encryption for all traffic
  2. Authentication: JWT with short expiration
  3. Authorization: Role-based access control (RBAC)
  4. Input Validation: Sanitize all user input
  5. Output Encoding: Prevent XSS attacks
  6. Database: Prepared statements prevent SQL injection

Data Privacy#

  • Encryption at rest: Depends on infrastructure provider's PostgreSQL configuration
  • Encryption in transit: TLS handled by reverse proxy (the application runs plain HTTP)
  • PII handling: PII lives in standard entity tables; logger-level redaction via REDACT_PATHS
  • Audit logs: Administrative actions (create/update/delete operations) are logged via AuditLogService

Monitoring and Observability#

Metrics#

Observability is currently limited to structured logging. The following metrics are planned but not yet collected (no metrics library is integrated):

  • API response times
  • Error rates
  • Queue lengths
  • Database query performance
  • Cache hit rates

Logging#

Structured logging with correlation IDs:

logger.info({ conversationId: '123', action: 'create', duration: 45 }, 'Processing conversation');

Tracing#

Distributed tracing is not yet implemented. Currently, error reporting is handled via opt-in PostHog exception capture (/server/lib/telemetry.ts). Full request tracing is planned for a future release.

Next Steps#

Last updated September 28, 2026

View on GitHub