Advertisement

Three Years of Looking at Data — Why Machine Learning Still Feels Like Magic (But Isn't)

Three Years of Looking at Data — Why Machine Learning Still Feels Like Magic (But Isn't)

On my 45-minute local train ride from Kalyan to Mumbai, I used to wonder why my Spotify playlist somehow knew I'd want to hear a Marathi folk song right after a lo-fi hip-hop track. I'd sit there, earbuds in, thinking: how does this app know me better than I know myself?

Then I started working in data analysis at Morningstar. And that's when the real confusion started.

Machine learning used to feel like black magic to me. Something only PhDs in Silicon Valley understood. Something that required code I'd never write and math I'd failed to grasp properly in college (Economics taught me a lot, but neural networks weren't exactly on the syllabus). But after three years of sitting across from data engineers, reading research papers on my commute, and honestly — breaking down how recommendation systems actually work — I've realized something: machine learning isn't magic. It's just pattern recognition on steroids. And more importantly, it's something every Indian working with data needs to understand, not because you'll build the next ChatGPT, but because it's reshaping how money moves, how careers progress, and how decisions get made in India.

Here's what I actually think about machine learning now, and why I'm writing this.

The Problem Nobody Admits: We're All Teaching Machines Now

Every time you use Zerodha to check your portfolio, or CRED to pay a credit card bill, or PhonePe to split rent with your roommate — you're feeding data into a machine learning system. You're a teacher. The machine is the student. And here's the uncomfortable truth: most of us have no idea what the student is learning.

When Groww shows you a mutual fund recommendation, that's not a human sitting somewhere. It's a model. A statistical pattern. Built from thousands of transactions like yours.

What Machine Learning Actually Is (Without the Jargon)

Let me strip this down to what matters.

Machine learning is a system that learns patterns from examples, not instructions. Traditional programming? You tell the computer: "If income is above ₹50 lakh and age is below 35, approve the loan." You write the rule. The computer follows it. Done.

Machine learning is different. You don't write the rule. You show the machine thousands of historical loan approvals. You say: "Here are 10,000 people. Some got loans. Some didn't. Figure out why." The machine then looks at age, income, credit score, employment history, savings patterns — everything. And it finds patterns you didn't explicitly tell it to find. Sometimes patterns you didn't even know existed.

That's it. That's machine learning in one paragraph.

The reason it feels like magic is because the rules it discovers are often invisible. You can't easily ask the algorithm "Why did you deny this person's loan?" In some cases, not even the engineers who built it can tell you precisely why. And that's genuinely scary when it involves someone's ₹25 lakh home loan application.

Why This Matters More in India Than You Think

India is unique because we're building these systems at scale but often without the infrastructure or regulation that Silicon Valley has. When a fintech app in Bangalore uses machine learning to decide your credit score, it might be using data points that seem unfair — like your location, or the apps on your phone, or your social media activity. In the US, there are strict rules against this. In India? We're still figuring it out.

This isn't a complaint. It's an observation. And it means understanding how these systems work isn't optional anymore. It's survival.

How Machine Learning Actually Works — The Honest Version

There are three main flavors of machine learning. And I'm going to explain them the way I wish someone had explained them to me three years ago, without making you feel stupid.

1. Supervised Learning — Teaching with a Clear Answer Key

Imagine you're teaching your little cousin multiplication. You show them: "2 + 2 = 4. 3 + 3 = 6. 4 + 4 = 8." Eventually, they learn the pattern. When you ask "5 + 5 = ?", they know.

That's supervised learning. You have historical data with known outcomes. You feed it into the system. The system learns. Then when new data arrives, it predicts.

Real example: Suppose a bank wants to predict who will default on a personal loan. They have 50,000 loan records from the past five years. For each record, they know if the person paid it back or defaulted. The machine learning model learns from these 50,000 examples. Now when a new ₹5 lakh loan application comes in, it can estimate the probability of default based on patterns it found in the historical data.

This is how most fintech apps in India decide whether to approve your loan instantly or make you wait for verification.

2. Unsupervised Learning — Finding Patterns Without Being Told

Now imagine you show someone a pile of 100 photographs without telling them anything about them. They naturally start sorting: "Oh, these are people. Those are landscapes. These look like they're from a temple." Nobody told them the categories. They found them.

That's unsupervised learning. You dump data into the system with no labels, no "correct" answers. The system finds patterns on its own.

Real example: Netflix uses this to group users into clusters. It doesn't know why you and I might be similar. But by analyzing what we watch, it discovers: "Okay, people in this cluster watch a lot of crime dramas, some international films, and old Hindi movies. Let's call this the 'Serious Viewer' cluster." Then it recommends shows based on what else people in your cluster watched.

This is genuinely useful. But it also means the algorithm is making assumptions about you that it never explicitly shares.

3. Reinforcement Learning — Learning by Trial and Error

Remember playing Snake on your old Nokia phone? You were learning through feedback: eat the food, score points, hit the wall, lose. Machine learning can work the same way. Show a system a goal, give it rewards for good actions and penalties for bad ones, and it learns to maximize reward.

Real example: Trading algorithms use this. A model learns to buy and sell stocks by simulating thousands of trading days. It gets rewarded when it makes profitable trades, penalized when it loses money. Over time, it learns a strategy.

This is also why you should be terrified of fully automated trading systems without human oversight.

Quick Tip: The type of machine learning used depends on the problem. Most fintech applications in India use supervised learning because we have tons of historical data to learn from. That's not coincidence. It's design.

The Actual Steps: From Raw Data to Your Loan Approval

Let me walk you through what actually happens when you apply for a ₹10 lakh personal loan on a fintech app.

Step 1: Data Collection. You fill out a form. Income. Age. Employment. Credit score. Assets. Liabilities. The system also pulls: your banking history from your connected bank account, your payment history on credit cards, even how quickly you fill out the form. All of this becomes data points.

Step 2: Data Cleaning. Some data is missing. Some is corrupted. Some is an outlier (like if you put your income as ₹100 crore on a form for a ₹10 lakh loan). Data scientists clean this. It's boring. It takes 60% of the time. Nobody talks about it because it's not sexy.

Step 3: Feature Engineering. Raw data isn't always useful. So engineers create new variables. Maybe it's not just "income" but "income relative to your city's average." Not just "age" but "age at first loan." These engineered features often matter more than raw data.

Step 4: Model Training. The system is shown historical loan data. It learns patterns. "People who took 15 days to repay their first ₹1 lakh loan are 3x more likely to default on bigger loans." "People in tier-2 cities with business income have a different risk profile than salaried employees in Mumbai." The model absorbs thousands of such patterns.

Step 5: Model Testing. Before using it on real applications, the model is tested on data it hasn't seen. If it predicts correctly 87% of the time on test data, it might be good enough to deploy. Or maybe not. Depends on the business cost of being wrong.

Step 6: Deployment and Monitoring. Your application hits the model. It spits out a probability of default. The app decides: approve or deny. But here's the critical part — the system is constantly monitored. If it's approving too many people who later default, it needs retraining. Machine learning isn't "set it and forget it." It's living, breathing, constantly evolving.

Stage What Happens Who's Involved
Data Collection You provide information via forms and APIs You + Product Team
Data Cleaning Remove missing/corrupted data, handle outliers Data Engineers
Feature Engineering Create new variables that might matter Data Scientists
Model Training System learns patterns from historical data ML Engineers, Data Scientists
Testing Validate on unseen data before deployment Data Scientists, QA
Deployment Model goes live and makes decisions on real data ML Engineers, DevOps
Monitoring Track accuracy, retrain as needed Analytics Team

Where Machine Learning Gets Messy — The Honest Conversation

I've been deliberately avoiding the uncomfortable parts. Time to address them.

Bias and Fairness — The Problem Nobody Has Solved

Here's a real scenario: A bank trains a machine learning model on 20 years of loan approval data. In that historical data, women and people from certain castes were approved for loans at lower rates. The model learns this pattern. When it goes live, it doesn't explicitly deny loans to women or marginalized groups — it's illegal. But it finds proxy variables. Maybe it's "preferred zip codes," which correlate with caste and wealth. Maybe it's "education level," which correlates with access and privilege.

The model isn't being deliberately discriminatory. But it's perpetuating historical discrimination. And here's the trap: the engineers who built it might not even notice.

In India, this is a massive issue that barely gets discussed. We have fintech companies making loan decisions for millions of people, and there's almost no regulatory oversight on algorithmic bias.

The Black Box Problem

You apply for a loan. It gets rejected. You ask why. The app says: "Sorry, you don't meet our criteria." You never find out if it was your income, your credit score, your location, or the fact that you searched for "loan default" on Google once.

Some machine learning models, especially deep learning models used for image recognition or natural language processing, are so complex that not even the engineers can explain their decisions. They call this the "black box problem." And it's legitimate.

When a machine learning model decides you're a credit risk, you deserve to know why. But in India, that's not guaranteed.

Data Privacy — The Elephant Nobody's Talking About

To build good machine learning models, you need lots of data. Which means companies are collecting everything about you: your bank statements, your browsing history, your social circle, your location. If that data leaks or gets misused, there's not much you can do.

I used to think privacy was a Western concern. Then I realized: if a machine learning model trained on my private financial data can be sold to competitors, or if my data gets breached, my entire financial life is exposed. That's not theoretical. It's happening.

Why Every Indian Millennial Should Care

Machine learning isn't just about fancy tech companies. It's directly affecting your career, your money, your opportunities.

Career Impact: If you work in any data-adjacent field — finance, consulting, product, operations — machine learning skills are becoming table stakes. Not "nice to have." Table stakes. I've watched junior analysts at my company get passed over for promotions because they couldn't engage with machine learning systems their teams were building. That stung to watch.

Investment Impact: When you buy stocks on Zerodha or invest in mutual funds via Groww, machine learning is deciding which companies you hear about, which funds are recommended to you. It's shaping your portfolio in ways you don't see.

Credit Impact: Your next personal loan, your credit card limit, your insurance premiums — all decided by machine learning. Understanding how it works gives you power. You know what kinds of data matter. You can optimize your behavior accordingly.

Social Impact: Hiring algorithms. Loan algorithms. Housing algorithms. These systems are making decisions that affect millions of Indians every day. If you understand them, you can at least spot when they're unfair.

My Perspective

Honestly, my commute from Kalyan to Mumbai has become thinking time about this. I'll be sitting in the local, watching people tap their phones to pay for tickets with PhonePe, and I'm wondering: what's the model deciding about them? What patterns has it found? Are those patterns fair?

I used to think machine learning was something only tech people needed to understand. I was wrong. It's reshaping how money moves in India, and if you're not thinking about it, someone else is deciding how your life works. That's the honest truth I've come to.

What surprised me: Machine learning isn't actually that complex once you strip away the jargon. It's just pattern matching. What terrifies me: we're rolling it out at scale in India without enough debate about fairness, privacy, and transparency. What I'd do differently: If I were starting fresh, I'd spend less time learning algorithms and more time understanding business context and ethics. The algorithms are tools. The real question is: should we be using them this way?

Final Thoughts

Machine learning isn't magic. It's pattern recognition. And you don't need to be a PhD to understand it. You just need to ask the right questions.

Why did this app deny my loan? What data is it using to decide? Could someone manipulate it? Who's responsible if something goes wrong?

These aren't technical questions. They're human questions. And they matter more than understanding backpropagation or gradient descent.

If you work with data, learn machine learning fundamentals. Not to become an ML engineer (unless you want to). But to be literate in a world where machines are making increasingly important decisions about you. Learn enough to ask good questions. Learn enough to spot bullshit. Learn enough to know when a company is hiding behind "it's AI" to avoid responsibility.

The future of India's financial system, its hiring systems, its credit systems — they're all being built on machine learning. You deserve to understand how they work. Not just for your career, but for your dignity.


Dattatray Dagale

Data Analyst • Blogger • Mumbai

I'm a data analyst from Kalyan, Maharashtra, working at Morningstar. I write about personal finance, career growth, and everyday life for Indian millennials — the stuff I wish someone had told me earlier.

Written by Dattatray Dagale • 03 October 2026

Post a Comment

0 Comments

×

📢 Featured Post

Post Thumbnail

💼 Budget 2025-26 💼

All major highlights.

📖 Read Now