I鈥檓 a Data Engineer with 6+ years of professional experience, currently working at Entual GmbH in Germany. My professional foundation is in Ab Initio and enterprise-scale data engineering, and my work has expanded into distributed systems, runtime behaviour, GPU computing, local AI, and performance-oriented software.
I鈥檓 especially interested in systems where data movement, execution models, memory, and hardware constraints directly shape performance and reliability.
Experimental runtime work focused on translating CUDA-oriented workloads toward Apple Metal, exploring runtime compatibility, execution models, and GPU portability on Apple Silicon.
C/C++ 路 CUDA 路 Metal 路 Apple Silicon 路 Runtime Systems 路 GPU Computing
A systems-focused project exploring local execution, tooling, and resource-aware compute workflows with an emphasis on practical infrastructure and performance-conscious design.
Rust 路 Systems Programming 路 Local Compute 路 Performance Engineering
My active development branch on llama.cpp, with work around persistent and file-backed KV cache, mmap-based memory optimisation, Vulkan/runtime improvements, GPU-backed cache behaviour, and server extensions.
C/C++ 路 LLM Inference 路 KV Cache 路 mmap 路 Vulkan 路 GPU Computing
Ab Initio and Data Engineering have been the core of my professional engineering career.
I have worked with enterprise-scale data-processing environments where correctness, reliability, scalability, recoverability, operational stability, and performance are critical.
Key areas include:
Ab Initio 路 GDE 路 Conduct>It 路 Control Center 路 Express>It 路 TRW
ETL / ELT 路 Data Integration 路 Data Quality 路 Batch Processing 路 Production Support 路 Performance Tuning
Broader data-engineering tools and technologies:
SQL 路 Oracle 路 PostgreSQL 路 Python 路 Shell / KornShell
I completed an M.Sc. in Data Science, which broadened that engineering foundation into machine learning, research workflows, and data-intensive systems.
My current interests sit at the intersection of data systems and lower-level compute:
- distributed and parallel data processing
- runtime and memory behaviour
- systems programming in Rust and C/C++
- GPU computing and hardware-aware optimisation
- local and resource-efficient AI
- LLM inference and cache architecture
- performance engineering across software and hardware boundaries
A recurring question in my work is: where is the real bottleneck, and what layer is actually responsible for it?
Data Engineering
Ab Initio 路 ETL / ELT 路 SQL 路 Oracle 路 PostgreSQL 路 Data Quality
Programming & Systems
Python 路 Rust 路 C/C++ 路 Shell 路 KornShell
Distributed & Parallel Systems
Kafka 路 Airflow 路 Ray 路 AsyncIO 路 Parallel Processing
Compute & AI Infrastructure
CUDA 路 Metal 路 Vulkan 路 Apple Silicon 路 LLM Inference 路 Local AI
Platforms
Linux 路 macOS
For a more complete view of my engineering work, projects, and background:
- Portfolio: https://perinban.github.io/portfolio/
- GitHub: https://github.com/Perinban
- LinkedIn: https://www.linkedin.com/in/perinban-parameshwaran/