A Futuristic Modular AI Operating System Built in Python
Features โข Quick Start โข Architecture โข Documentation โข Contributing
- Overview
- Features
- Tech Stack
- Project Structure
- Prerequisites
- Installation
- Environment Setup
- Quick Start
- Usage
- API Integration
- Configuration
- Screenshots
- Development
- Roadmap
- Contributing
- Troubleshooting
- License
- Author
ARAXON is an advanced modular AI operating system that integrates multiple AI capabilities into a cohesive, scalable platform. It combines voice interaction, computer vision, persistent memory, automation capabilities, and autonomous agents into a single, extensible framework.
Whether you're building an AI assistant, automating workflows, or exploring advanced AI capabilities, ARAXON provides a production-ready foundation with professional-grade architecture.
- ๐๏ธ Voice-First Interaction: Natural language processing with speech recognition and synthesis
- ๐๏ธ Computer Vision: Advanced image processing and object detection
- ๐ง Persistent Memory: Vector-based knowledge base with semantic search
- ๐ค Autonomous Agents: Multi-agent system with advanced decision making
- ๐ Automation: Workflow orchestration and task automation
- ๐ Internet Integration: Web browsing, API access, and data extraction
- โก Async-First: Built on Python asyncio for high-performance concurrent operations
- ๐จ Beautiful UI: Modern React-based interface with Tauri desktop integration
- ๐๏ธ Speech-to-text with Faster-Whisper
- ๐ Text-to-speech synthesis
- ๐ Wake word detection and voice activation
- ๐ฏ Natural language understanding
- ๐ธ Real-time screenshot capture
- ๐ Optical Character Recognition (OCR)
- ๐๏ธ Computer vision analysis
- ๐ฌ Video processing and streaming
- ๐ Vector database with ChromaDB
- ๐ง Semantic search capabilities
- ๐พ Long-term knowledge storage
- ๐ Document ingestion and processing
- ๐ Knowledge graph integration
- ๐ค Multi-agent orchestration
- ๐ Task automation workflows
- โฑ๏ธ Intelligent scheduling
- ๐ฏ Goal-oriented execution
- ๐ Agent performance monitoring
- ๐ Web browsing automation
- ๐ API integration
- ๐ฐ News aggregation
- ๐ Intelligent web search
- ๐ฐ Content extraction and analysis
- ๐จ Modern React dashboard
- ๐ป Tauri desktop application
- ๐ WebSocket real-time communication
- ๐ฑ Responsive design
- ๐ฏ Intuitive command interface
- Python 3.11+ - Core runtime
- LangChain - AI orchestration framework
- LangGraph - Agent and workflow execution
- Groq & Ollama - LLM providers
- ChromaDB - Vector database
- Faster-Whisper - Speech recognition
- Kokoro - Text-to-speech synthesis
- PyTorch - Deep learning framework
- Sentence Transformers - Embedding models
- OpenCV - Computer vision
- Tesseract - OCR engine
- NumPy - Numerical computing
- React 18+ - UI framework
- Tauri - Desktop application
- Vite - Build tool
- WebSockets - Real-time communication
- CSS3 - Styling
- asyncio - Asynchronous programming
- Pydantic - Data validation
- Loguru - Logging framework
- python-dotenv - Environment management
ARAXON/
โโโ araxon/ # Main package
โ โโโ __init__.py
โ โโโ core/ # Core infrastructure
โ โ โโโ config.py # Configuration management
โ โ โโโ logger.py # Logging system
โ โ โโโ utils.py # Utility functions
โ โ
โ โโโ ai/ # AI integration
โ โ โโโ brain.py # Core AI reasoning
โ โ โโโ router.py # Request routing
โ โ โโโ memory.py # AI memory management
โ โ โโโ personality.py # AI personality traits
โ โ
โ โโโ voice/ # Voice I/O
โ โ โโโ listener.py # Speech input
โ โ โโโ synthesizer.py # Speech output
โ โ โโโ transcriber.py # Audio transcription
โ โ โโโ audio_player.py # Audio playback
โ โ โโโ voice_*.py # Voice utilities
โ โ
โ โโโ vision/ # Computer vision
โ โ โโโ analyzer.py # Image analysis
โ โ โโโ screenshot.py # Screenshot capture
โ โ โโโ ocr.py # Optical character recognition
โ โ โโโ vision_pipeline.py # Processing pipeline
โ โ โโโ vision_router.py # Vision routing
โ โ
โ โโโ memory/ # Knowledge & memory system
โ โ โโโ embedder.py # Embedding generation
โ โ โโโ vector_store.py # Vector database wrapper
โ โ โโโ long_term_memory.py # Persistent memory
โ โ โโโ file_ingester.py # Document processing
โ โ โโโ rag_pipeline.py # RAG implementation
โ โ
โ โโโ automation/ # Task automation
โ โ โโโ automation_router.py # Request routing
โ โ โโโ app_launcher.py # Application launching
โ โ โโโ command_runner.py # Command execution
โ โ โโโ browser_agent.py # Browser automation
โ โ โโโ workspace_manager.py # Workspace management
โ โ
โ โโโ agent/ # Agent system
โ โ โโโ agent_controller.py # Agent orchestration
โ โ โโโ executor.py # Task execution
โ โ โโโ planner.py # Planning & reasoning
โ โ โโโ graph.py # Agent graph
โ โ โโโ tools.py # Agent tools
โ โ
โ โโโ internet/ # Internet integration
โ โ โโโ internet_router.py # Request routing
โ โ โโโ searcher.py # Web search
โ โ โโโ news_fetcher.py # News aggregation
โ โ โโโ researcher.py # Research capabilities
โ โ โโโ wiki_lookup.py # Wikipedia integration
โ โ โโโ extractor.py # Content extraction
โ โ
โ โโโ ui/ # User interface
โ โโโ ui_bridge.py # Frontend bridge
โ โโโ websocket_server.py # Real-time communication
โ
โโโ ui/ # Frontend application
โ โโโ src/ # React source
โ โ โโโ components/ # UI components
โ โ โโโ App.jsx # Main app
โ โ โโโ main.jsx # Entry point
โ โโโ src-tauri/ # Tauri backend
โ โโโ package.json # Frontend dependencies
โ โโโ vite.config.js # Build configuration
โ
โโโ config/ # Configuration
โ โโโ settings.yaml # Settings file
โ
โโโ data/ # Data storage
โ โโโ chromadb/ # Vector database
โ โโโ ingested/ # Processed documents
โ
โโโ logs/ # Application logs
โ โโโ screenshots/ # Captured screenshots
โ
โโโ models/ # AI models
โ โโโ models--* # Model cache
โ
โโโ main.py # Application entry point
โโโ requirements.txt # Python dependencies
โโโ setup_project.py # Project initialization
โโโ .env.example # Environment template
โโโ README.md # Documentation
โโโ LICENSE # License file
โโโ .gitignore # Git ignore rules
Before you begin, ensure you have the following installed:
- OS: Windows 10/11, macOS 11+, or Linux (Ubuntu 20.04+)
- Python: 3.11 or higher
- Node.js: 16+ (for frontend development)
- RAM: Minimum 8GB (16GB recommended for ML models)
- GPU: NVIDIA GPU recommended (CUDA 11.8+) for faster inference
- Git
- pip (Python package manager)
- FFmpeg (for audio processing)
- Tesseract OCR (for document processing)
- Ollama (for local LLM inference)
- Groq API key (for cloud LLM access)
- Chromium/Chrome (for browser automation)
git clone https://github.com/yourusername/araxon.git
cd araxonWindows:
python -m venv .venv311
.venv311\Scripts\activatemacOS / Linux:
python3 -m venv .venv311
source .venv311/bin/activatepip install --upgrade pip setuptools wheel
pip install -r requirements.txtWindows (PowerShell as Admin):
# Install FFmpeg
choco install ffmpeg -y
# Install Tesseract OCR
choco install tesseract -y
# Optional: Install Ollama
choco install ollama -ymacOS:
# Install FFmpeg
brew install ffmpeg
# Install Tesseract OCR
brew install tesseract
# Optional: Install Ollama
brew install ollamaLinux (Ubuntu/Debian):
sudo apt-get update
sudo apt-get install -y ffmpeg tesseract-ocr
# Optional: Install Ollama (https://ollama.ai)If you want to develop the UI:
cd ui
npm install
npm run build # or npm run dev for developmentCopy the provided .env.example file:
cp .env.example .env# Application Settings
APP_NAME=ARAXON
DEBUG_MODE=False
LOG_LEVEL=INFO
# API Keys
GROQ_API_KEY=your_groq_api_key_here
OPENAI_API_KEY=your_openai_api_key_here # Optional
# LLM Configuration
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=mistral # or llama2, neural-chat, etc.
# Voice Configuration
WAKE_WORD=araxon
TTS_MODEL=kokoro # or tts-1, etc.
SPEECH_RECOGNITION_LANGUAGE=en-US
# Memory Configuration
CHROMA_DB_PATH=./data/chromadb
VECTOR_STORE_TYPE=chroma
# Database
DATABASE_URL=sqlite:///./data/araxon.db
# UI Configuration
UI_HOST=localhost
UI_PORT=5173
WEBSOCKET_PORT=8765
# Optional: LLM Provider Selection
LLM_PROVIDER=groq # options: groq, ollama, openai
# Optional: Vision Configuration
VISION_MODEL=gpt-4-vision # or local model
OCR_LANGUAGE=engGroq API Key:
- Visit console.groq.com
- Sign up for a free account
- Create an API key
- Add to your
.envfile
OpenAI API Key (Optional):
- Visit platform.openai.com
- Create an account
- Generate API key
- Add to your
.envfile
# Activate virtual environment first
.venv311\Scripts\activate # Windows
# or
source .venv311/bin/activate # macOS/Linux
# Run setup script
python setup_project.py
# Start ARAXON
python main.pyTerminal 1 - Backend:
.venv311\Scripts\python.exe main.pyTerminal 2 - Frontend (Optional):
cd ui
npm run devThe application will:
- Initialize all subsystems
- Load configuration from
.env - Connect to vector database
- Initialize voice and vision pipelines
- Start the web server
- Begin listening for voice commands
python main.pyThen interact with ARAXON:
- Voice: Speak your wake word followed by a command
- CLI: Type commands directly in the terminal
- Web UI: Access the dashboard at
http://localhost:5173
# Voice Command
"Araxon, search for Python tutorials"
"Araxon, take a screenshot"
"Araxon, what time is it?"
# Terminal Command
python main.py --command "search python tutorials"
python main.py --voice-onlyimport asyncio
from araxon.ai import ARAXONBrain
from araxon.agent import AgentController
async def main():
# Initialize brain
brain = ARAXONBrain()
# Process a query
response = await brain.process("What is AI?")
print(response)
# Interact with agents
agent = AgentController()
result = await agent.execute_task("Search for recent AI news")
print(result)
asyncio.run(main())ARAXON exposes several REST APIs:
POST /api/brain/think
Content-Type: application/json
{
"query": "What is machine learning?",
"context": "optional_context"
}
Response: { "response": "...", "confidence": 0.95 }
POST /api/vision/analyze
Content-Type: multipart/form-data
File: image.png
Response: { "analysis": "...", "objects": [...] }
POST /api/memory/query
Content-Type: application/json
{
"query": "Find documents about Python",
"limit": 10
}
Response: { "results": [...], "total": 42 }
POST /api/agent/execute
Content-Type: application/json
{
"task": "Automate daily report generation",
"parameters": {...}
}
Response: { "status": "completed", "result": "..." }
const ws = new WebSocket('ws://localhost:8765');
ws.onmessage = (event) => {
console.log('Message from ARAXON:', event.data);
};
ws.send(JSON.stringify({
type: 'command',
data: 'your command here'
}));from pydantic_settings import BaseSettings
class Settings(BaseSettings):
APP_NAME: str = "ARAXON"
DEBUG_MODE: bool = False
LOG_LEVEL: str = "INFO"
# LLM Settings
LLM_PROVIDER: str = "groq"
GROQ_API_KEY: str
class Config:
env_file = ".env"
case_sensitive = Truefrom araxon.core.logger import logger
logger.info("Application started")
logger.debug("Debug information")
logger.error("An error occurred")
logger.warning("Warning message")Edit config/settings.yaml:
application:
name: ARAXON
version: 1.0.0
ai:
model: mistral
temperature: 0.7
voice:
wake_word: araxon
language: en-US
memory:
vector_dimension: 1536
top_k: 10[Screenshot placeholder - Update with actual dashboard screenshot]
[Screenshot placeholder - Update with actual UI screenshot]
[Screenshot placeholder - Update with actual vision output screenshot]
[Screenshot placeholder - Update with actual agent output screenshot]
๐ก Tip: Replace these placeholders with actual screenshots of your application
ARAXON follows these principles:
- Modularity: Each subsystem is independent
- Async-First: All I/O operations use async/await
- Configuration: Settings via environment variables
- Logging: Comprehensive logging with loguru
- Testing: Dedicated test files for each module
- Create a new folder in
araxon/ - Add
__init__.pywith module exports - Implement your module with async functions
- Register in
main.pyinitialization - Add tests in
tests/
Example:
# araxon/my_module/__init__.py
from .my_module import MyModuleClass
__all__ = ["MyModuleClass"]
# Usage in main.py
from araxon.my_module import MyModuleClass
async def initialize():
my_module = MyModuleClass()
await my_module.setup()# Run all tests
pytest
# Run specific test
pytest tests/test_core.py
# Run with coverage
pytest --cov=araxon tests/ARAXON uses:
- Black for formatting
- Flake8 for linting
- MyPy for type checking
# Format code
black araxon/
# Lint
flake8 araxon/
# Type check
mypy araxon/- Foundation infrastructure
- Configuration management
- Logging system
- Core utilities
- Advanced voice capabilities
- Vision pipeline optimization
- Memory system enhancement
- UI/UX improvements
- Multi-language support
- Enhanced security features
- Mobile app version
- Cloud synchronization
- Plugin system
- Performance optimization
- Full voice AI assistant
- Advanced reasoning capabilities
- Custom model training
- Enterprise features
- API marketplace
See PROJECT_CHECKLIST.md for detailed progress.
We welcome contributions! Here's how to help:
Please note we have a Code of Conduct to ensure a welcoming community.
- Fork the repository
- Create a feature branch:
git checkout -b feature/amazing-feature - Make your changes
- Write tests for new functionality
- Commit:
git commit -m 'Add amazing feature' - Push:
git push origin feature/amazing-feature - Open a Pull Request
- Follow PEP 8 style guide
- Write docstrings for all functions
- Add type hints
- Include unit tests (minimum 80% coverage)
- Update documentation
- Use conventional commit messages
<type>(<scope>): <subject>
<body>
<footer>
Types: feat, fix, docs, style, refactor, test, chore
Example:
feat(vision): add image recognition capability
Implemented ResNet-based image classification with
confidence scoring and bounding box detection.
Closes #123
- Update
README.mdif needed - Update
CHANGELOG.mdwith changes - Ensure tests pass:
pytest - Request review from maintainers
- Address feedback and iterate
Solution:
# Ensure virtual environment is activated
.venv311\Scripts\activate # Windows
source .venv311/bin/activate # macOS/Linux
# Reinstall dependencies
pip install -r requirements.txtSolution:
# Check .env file exists and is readable
cat .env # or type .env on Windows
# Get API key from https://console.groq.com
# Add to .env:
GROQ_API_KEY=your_key_hereSolution:
# Check device permissions
# Verify audio device:
python -c "import sounddevice; print(sounddevice.query_devices())"
# Reinstall audio libraries
pip install --upgrade sounddevice numpySolution:
# Check CUDA installation
python -c "import torch; print(torch.cuda.is_available())"
# Install CUDA-enabled PyTorch
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118Solution:
# Reset vector database
rm -rf data/chromadb/*
# Reinitialize
python -c "from araxon.memory import LongTermMemory; m = LongTermMemory(); m.initialize()"Enable detailed logging:
# Via command line
DEBUG_MODE=True LOG_LEVEL=DEBUG python main.py
# Or in .env
DEBUG_MODE=True
LOG_LEVEL=DEBUG- ๐ Check INSTALLATION.md
- ๐ Review DEVELOPER_GUIDE.md
- ๐ฌ Open an Issue
- ๐ Contact the maintainers
This project is licensed under the MIT License - see the LICENSE file for details.
The MIT License is a permissive open-source license that allows you to:
- โ Use commercially
- โ Modify the software
- โ Distribute the software
- โ Use privately
- โ Hold liable
With the requirement:
- ๐ Include license and copyright notice
For more information, visit opensource.org/licenses/MIT
ARAXON Development Team
- GitHub: @yourusername
- Email: your.email@example.com
- Website: your-website.com
- Contributor 1 - Feature/Module contributions
- Contributor 2 - Bug fixes and improvements
Special thanks to:
- The Python community
- LangChain team for excellent framework
- All contributors and supporters
- Open-source projects we depend on
- ๐ Report Bugs: GitHub Issues
- ๐ก Feature Requests: GitHub Discussions
- ๐ Documentation: Wiki
- ๐ฌ Community Chat: Discord
- ๐ค Join Us: Contributing Guide
- Language: Python 3.11+
- License: MIT
- Repository: GitHub
- Status: Active Development ๐
- Last Updated: May 2024
For security concerns, please email: security@your-domain.com
Do not open security vulnerabilities publicly. Please follow responsible disclosure practices.
See CHANGELOG.md for detailed version history.
- Version: 1.0.0
- Released: May 2024
- Release Notes
Made with โค๏ธ by the ARAXON Team
โญ Star us on GitHub | ๐ฆ Follow us on Twitter | ๐ง Newsletter
Last Updated: May 24, 2024
Documentation Version: 1.0.0
Status: โ
Maintained