I am a
Building production-ready ML systems and data-driven insights.
Designing intelligent systems that transform complex data into scalable, real-world solutions. From machine learning pipelines and MLOps to generative AI and analytics, every project is engineered with precision, optimized for performance, and built to deliver measurable business impact.
About My Expertise
As a Data Scientist and Machine Learning Engineer, I specialize in architecting end-to-end MLOps, scalable data pipelines, and production-ready AI systems. My core focus is bridging the gap between cutting-edge research and measurable business value through rigorous experimentation and modern AI technologies.
I have a proven track record of delivering intelligent data-driven insights and scalable ML solutions across domains, from predictive analytics to Natural Language Processing (NLP). Whether it's fine-tuning Large Language Models (LLMs) or optimizing data engineering workflows, I build systems designed for reliability and performance.
MLOps & Production ML
Deployment, monitoring, and CI/CD for machine learning models using Docker, Kubernetes, and cloud native tools.
AI & RAG Applications
Building advanced RAG systems, custom LLM agents, and conversational AI tailored to specific data needs.
Predictive Analytics
Transforming complex datasets into predictive insights and actionable business intelligence dashboards.
Data Engineering
Architecting robust ETL pipelines, SQL optimizations, and real-time data processing for enterprise scale.
Skills & Technologies
Tools and technologies I work with every day
Currently Exploring
Areas I'm actively studying and building expertise in right now.
Agentic AI & Multi-Agent Systems
Building autonomous AI agents with tool use, planning, and memory using LangGraph and CrewAI frameworks.
Kubernetes for MLOps
Deploying and scaling ML workloads on Kubernetes clusters using Kubeflow and custom operators.
LLM Fine-Tuning (QLoRA / PEFT)
Fine-tuning large language models efficiently using Parameter-Efficient Fine-Tuning techniques on domain-specific datasets.
Rust for Systems Programming
Learning Rust to build high-performance data processing tools and extend Python ML pipelines with native extensions.
Graph Neural Networks
Studying GNNs for fraud detection, recommendation systems, and knowledge graph construction using PyG.
Data Contracts & Data Quality
Implementing data contracts, schema validation, and automated quality checks in production data pipelines.
Featured Project
Financial Document Analysis Dashboard
An AI-powered Streamlit dashboard that extracts, parses, and analyzes financial documents (PDFs, reports, statements) to provide structured insights and visualizations for decision-making. Features intelligent NLP parsing, regex-based entity extraction, and interactive chart visualizations.
Featured Projects
A selection of my work in ML, AI, and data engineering
GitHub Activity
Monitoring real-time contribution grids and exploring active repositories directly from my GitHub profile.
@avijit-jana
contributions in
Certifications
Professional certifications and credentials that validate my expertise.
Machine Learning Engineering for Production (MLOps)
DeepLearning.AI
Machine Learning Specialization
Stanford University / Coursera
AWS Certified Cloud Practitioner
Amazon Web Services
Python for Data Science and AI
IBM / Coursera
Google Data Analytics Professional Certificate
Google / Coursera
LangChain & Vector Databases in Production
Activeloop
Latest from the Blog
Real-time AI research from the world's leading labs
Let's Connect
I'm available for collaboration, consulting, or tackling data science challenges. Whether it's a project, an idea, or a new opportunity — I'd love to hear from you.
Data Science & AI Solutions FAQ
Insights into engineering methodology, model optimization techniques, and business impact strategy.
What services do you provide as a Data Scientist & AI Engineer?
Specializing in end-to-end machine learning engineering, MLOps implementation, and custom Generative AI systems. This spans data architecture, model development, and deploying scalable inference pipelines on AWS or GCP.
How do you approach an enterprise Machine Learning lifecycle?
Following a structured MLOps workflow: aligning business objectives, performing rigorous exploratory data analysis (EDA), training baseline models, and deploying containerized solutions equipped with automated data drift monitoring.
How do you build production-ready RAG applications for enterprise data?
Building scalable Retrieval-Augmented Generation systems using vector stores like Pinecone or Qdrant, implementing hybrid search (dense + sparse) and re-ranking algorithms to eliminate hallucinations and secure query responses.
How do you decide between Fine-Tuning an LLM and using RAG?
Selected based on the use case: RAG is optimal for dynamic, external knowledge retrieval with strict data freshness, whereas Fine-Tuning is chosen when modifying a model's specialized style, tone, or domain-specific grammar.
How do you optimize deep learning models for fast, low-cost inference?
Utilizing model quantization (INT8/FP16), structured pruning, and hardware acceleration frameworks like ONNX Runtime or TensorRT to reduce memory footprint and latency while retaining predictive precision.
How do you evaluate and demonstrate real business ROI for AI projects?
Defining baseline metrics before development, optimizing cost-per-inference ratios, and executing A/B testing frameworks to measure direct business impacts on conversion rates, automation throughput, and operational costs.
