Today I sat down to write this with one very specific intention: I am not going to write anything in difficult language. I already have a bit of a bad reputation for this. Whenever I start writing about technology, I somehow manage to make it complicated. There is actually a reason for that, to be honest. Whenever I write about something technological, I keep in mind that the piece will be read by people who know very little about technology as well as by the techy crowd. If I use too much jargon, the non-techies will get annoyed. But if I simplify things too much, the tech nerds will get annoyed instead. So today, I have decided to annoy the techies by oversimplifying the whole thing.
Anyway, let me tell you what I actually sat down to write about. Today, we are going to talk about something rather magical.
The magic of technology.
Let me begin by giving you two paragraphs. Read them carefully.
First Paragraph: পুরো দরবার একদম নিস্তব্ধ। কারোর মুখে যেন কথা ফুটছে না। কারণ চোখের সামনে যা ঘটলো, তা বিশ্বাস করাও কঠিন। কিছুক্ষণ আগেও যেখানে রাজ্যের বড় বড় বীরেরা ব্যর্থ হয়ে ফিরেছে এবং বিশাল ওই ধনুকটিতে সামান্য টান পর্যন্ত দিতে পারেনি, সেখানে এক অচেনা প্রবীণ লোক এত সহজে এই অসম্ভবকে সম্ভব করে ফেলল! চোখের সামনে ঘটে যাওয়া এই দৃশ্য দেখে উপস্থিত সকলের মাথায় যেন বাজ পড়ল। তবে কি শেষ পর্যন্ত পেনেলোপির স্বামী হওয়ার যোগ্যতা অর্জন করে নিল এক হতদরিদ্র পথিক? ঠিক পরক্ষণেই পুরো দরবারে একটি সন্দেহের ঢেউ খেলে গেল—বহু বছর আগে যুদ্ধক্ষেত্রে হারিয়ে যাওয়া স্বয়ং ওডিসাসই কি তবে বেশ বদলে ফিরে এলেন?
Second Paragraph: প্রাসাদের বিরাট হলরুমটি একদম শান্ত হয়ে গেল। কারণ এইমাত্রই সবাই বিস্ময়ে প্রত্যক্ষ করলো যে, এক জীর্ণশীর্ণ বৃদ্ধ বীর ইতিফাসের সেই বিরাট ধনুকটিকে বাঁকিয়ে সেটায় ছিলা পরিয়ে পরপর ১২টি কুঠারের ছিদ্রের মধ্য দিয়ে তীর পাঠালেন! এমনটা কেউ-ই আশা করেনি। রাজ্যের কত কত বীরপুরুষ সে চেষ্টা করলেন—কিন্তু ধনুকটিকে বাঁকানোই সম্ভব হলো না তাঁদের পক্ষে। কিন্তু এই হতদরিদ্র পথিক যেন কোনো দক্ষ সুরকারের মতো বীণার তারে সুর বাঁধার মতো ধনুকটি দিয়ে তীর ছুঁড়লেন। তবে কি পেনেলোপিকে পেতে যাচ্ছে এই জীর্ণশীর্ণ ভিখারি? পরক্ষণেই সবার মনে পড়ল, এই ভিখারিই কি ওডিসাস, যে ১০ বছর আগে যুদ্ধে গিয়ে আর ফিরে আসেননি?
(I am sorry if you are not a native bengali speaker. I am intentionally keeping these two paragraphs in Bengali because they are the original texts being compared in the article.)
Now, tell me this: which of the two paragraphs above was written by a human, and which one was written by AI?
I’ll give you the answer later. But you have to admit one thing: you do not know with complete certainty which one was written by a human and which one was written by AI. In fact, you probably cannot even tell whether I wrote this article myself or had an AI polish it for me.
So, what’s the solution?
Does that mean we could simply have AI write our research papers, books, and everything else for us and go on to make a huge name for ourselves in the world?
No. We couldn’t.
Because even if you and I cannot tell whether a piece of writing was produced by AI or by a human, there is a little magic called “machine learning” that can help us figure it out. So how do we build an AI-generated text detector that can tell us whether a piece of writing was produced by AI?
Before that, let’s understand, in the simplest possible way, how AI itself can produce meaningful stuff, line after line.
Machines are essentially taught to write through statistics and mathematical probability. You can see a tiny version of this right on your keyboard. Whatever word you tend to use after another word, your keyboard learns that pattern and suggests it to you. With a Large Language Model, or LLM, however, the idea works on a much, much larger scale. An AI model has learned from billions of pieces of writing how to place one sentence after another to produce something meaningful.
Imagine training it on several thousand Bengali books. In those books, poets and writers might write things like, “Today, the sky is filled with clouds,” “The sky is covered with clouds,” or “The sky over Dhaka will remain cloudy throughout the day.” When a model repeatedly encounters patterns like these, it gradually learns to predict the probability of one word appearing after another.
It learns that if you write, “Today the sky is full of…” the probability of the next word being “clouds” is very high, while the probability of it being “Donald Trump” is almost zero. So, every time it generates text, it chooses from the possible words and keeps picking the one with the highest probability, eventually building the entire piece of writing.

If I put on my nerd hat for a moment, I would say that after AI writes one word, it calculates a probability distribution for what the next possible word could be. In mathematical terms, AI-generated writing generally tends to occupy the “local maxima” of a certain low-probability function, meaning the points with the highest probability. In other words, AI tends to choose the safest and most probable word available.
I don’t know how much of that you understood, but hopefully you at least understand this much: the process by which AI generates writing is bound by a mathematical framework of probabilities. And that is exactly where we can strike.
Since AI cannot simply step outside its underlying mathematical structure and write however it wants, we need to recognize that mathematical structure itself. If we can identify it, we can identify which writing was produced by AI and which was not. But as humans, we may not be able to do that ourselves.
So what do we do?
The answer is pretty simple.
We build another AI.
We can build another AI that can detect what the mathematical structure and statistical signature of AI-generated writing look like.
For that, we would need thousands of text samples written by AI and thousands of text samples written by humans. After collecting them, we would label them to indicate which texts are AI-generated and which are human-written.
Then we would choose an “encoder” that knows how to read text. Ahh! There I go throwing jargon at you again! We definitely do not want you to come across some incomprehensible technical term and leave the article halfway through.
So, what is an encoder?
An encoder is a special machine, model, or process. When you give it some text, instead of storing that text simply as its meaning, it converts it into a vector representation. This representation contains information about the words in the text, its context, and the relationships between different words.
Suppose you give an encoder the sentence “আমি ভাত খাই” (“I eat rice”). It would turn it into something roughly like this:

So, the BUET NLP team has built a wonderful encoder called BanglaBERT. For classification tasks in the Bengali language, I personally consider this encoder to be the best I have come across (personal opinion). This encoder does not simply look at each word separately. Instead, it tries to understand the context of the words within a sentence and the relationships between them. So that makes our job a little easier, doesn’t it?
We can then grab this BanglaBERT encoder by the collar and teach it, explain to it, how AI writes and how AI “does not write.”
Remember that huge collection of AI-written and human-written texts I mentioned earlier? Let’s feed those into this transformer encoder. It reads the texts and makes its own prediction: “Okay, is this written by AI or by a human?” It will make mistakes, and we will correct it. This cycle will continue hundreds of thousands of times. We will never explicitly tell it, “Listen, you idiot, if you see too many em dashes ‘—’, too many punctuation marks, perfect grammar, or non-ASCII symbols, flag it as AI.” It will figure out for itself which characteristics distinguish human writing from AI writing.

After some time, with proper data preprocessing and proper training, your model can genuinely learn those subtle statistical patterns that exist between AI-generated and human-written text. Then, when you give it a completely new piece of writing, it can make a fairly good prediction about whether that text is more likely to have been written by AI or by a human.
Don’t we, in real life, arrive at something similar ourselves? Suppose you are an experienced jeweler and someone puts a fake diamond and a real diamond in front of you. You can easily identify which one is real and which one is fake. Did you develop that level of mastery in a single day?
Didn’t you train your brain to recognize genuine gems by handling all kinds of gemstones hundreds of times over hundreds of days?
The world of teaching machines, or machine learning, works in a somewhat similar way. And sometimes that learning can become even more precise than our own. We can make a machine detect extremely subtle patterns that we ourselves would never be able to notice. This is why our fear that AI may one day take away our jobs is not entirely unfounded. In fact, even in something as sensitive as medicine, AI models are already achieving greater performance than humans in certain areas. Just last year, in 2025, I came across a research paper. A model developed by Google called AMIE was 26% ahead of physicians in Top 10 accuracy across 302 challenging real-world medical cases.
So, when AI can recognize all sorts of things with such remarkable precision, and can distinguish between two pieces of writing that look similar by telling us which one was written by AI and which one was written by a human, what exactly is it looking at? Is it seeing some kind of “mark” or watermark in AI-generated writing or human writing? None of us really knows. We know how AI learns, but we do not know exactly what it “learns.” Even AI godfather Geoffrey Hinton himself is quite concerned about this black box. Anyway, I don’t want to scare you so much that I end up hurting my own livelihood. This is how I make my living, after all.
Now, let me give you the results of the two paragraphs I gave you at the beginning. The second one is mine, meaning it was written by a “human.” And the first one was written by Gemini, meaning it was written by “AI.”
Want to be even more certain? Head over to the BanglaTuring we built: https://banglaturing.pages.dev/
Yes, we conducted research on AI detection in the Bengali language for our AI-powered fact-checking platform, “Khoj”, and we even deployed a model for it. Go ahead and verify it yourself. Although it is a little less capable when it comes to paraphrased text, it can still correctly identify whether a piece of writing was generated by AI or written by a human.
How did we build BanglaTuring? That magical story is for another day. But that day, I won’t make the story quite this simple. There will be all sorts of spells and sorcery appearing throughout the story.
Until then, stay hungry, stay foolish…






মতামত জানান