Exploring the Fundamentals of Natural Language Processing: An NLP Mid-Term Insight
Natural Language Processing connects human language with machine understanding. Source: Akamai Technologies.Natural Language Processing, or NLP, sits at an interesting intersection between two worlds: human language and computation.
This article grew out of my mid-term examination for the Natural Language Processing course. The exam itself was fairly conventional, including handwritten answers. What interested me more, however, was how several fundamental NLP questions made me look at language from a different perspective.
Coming from a literature background, I had previously encountered language through meaning, structure, context, and interpretation. In NLP, many of those same elements appear again, but with a different question:
How can human language be represented in a form that machines can understand and process?
The two questions in this exam offered a simple way to explore that problem.
Question 1
The first question asked us to explain the main differences between formal language, such as programming languages or logical systems, and natural language, such as English.
The most fundamental difference lies in how they are designed and how they develop.
Formal languages are created for specific purposes. Programming languages, formal logic, mathematical notation, and regular expressions require clear rules so they can be processed consistently. Their syntax and semantics are designed to minimise ambiguity.
Natural language works very differently.
Human languages develop through use, culture, social interaction, and change over time. Because of that, their structures and meanings are much more flexible. A word can have several meanings, a sentence can be interpreted differently depending on context, and idioms often cannot be understood simply by reading each word literally.
This is where NLP becomes particularly interesting to me.
In literature, ambiguity is not necessarily a problem. A word or sentence that allows multiple interpretations can even contribute to the richness of a text.
For a computer, however, ambiguity becomes a challenge.
Humans can rely on context, prior knowledge, experience, or even tone to understand what someone means. Machines need mechanisms that allow them to approximate that process.
This is one reason natural language is much more difficult to process automatically than formal language.
Formal language prioritises precision and consistency, while natural language brings something much more complex: context, ambiguity, flexibility, and meaning.
For me, this first question showed that one of the central challenges of NLP comes from the nature of human language itself.
Question 2
The second question took the problem one step further.
We were asked to imagine an NLP system that could receive premises in English and perform simple reasoning. The example used the classic syllogism about Socrates:
All men are mortal. Socrates is a man. Is Socrates mortal?
For a human reader, the conclusion seems immediate:
Yes.
A computer, however, does not simply read those sentences in the way a human does. Several components are needed to transform natural language into information that can be processed logically.
The process begins with Natural Language Understanding (NLU). A tokenizer breaks the sentence into units that can be processed, a parser analyses its grammatical structure, an entity recogniser identifies entities such as Socrates, and the system needs to distinguish between statements and questions.
The natural language input then needs to be transformed into a logical representation.
The sentence:
All men are mortal
can be represented as:
∀x (man(x) → mortal(x))
This means that for every x, if x is a man, then x is mortal.
The second statement:
Socrates is a man
can be represented as:
man(Socrates)
These representations can then be stored in a knowledge base, where the system keeps the facts and rules it has received.
When the user asks whether Socrates is mortal, an inference engine can apply the available rule and fact to derive:
mortal(Socrates)
The logical result can then be converted back through response generation into something easier for a human to understand:
Yes, Socrates is mortal.
The process can be simplified as:
human language → understanding → logical representation → knowledge base → inference → human language
What I find interesting is how many steps are required for something that feels almost automatic to us.
A human reads the two premises and understands the relationship between them. A machine needs an explicit representation before it can reach the same conclusion.
What I Learned from the Exam
The two questions may seem like basic NLP material, but they made the relationship between language and computation much clearer to me.
Language that I had previously encountered through literature appeared again in a different form. Words can become tokens.
Sentence structure can be analysed through parsing.
Meaning and relationships between concepts need to be transformed into logical representations.
Something that humans understand through context needs to be made more explicit before a machine can perform inference.
The terminology is different, and the tools are different, but the object being studied remains the same: language.
On one side, language can be studied as a medium through which humans express ideas, experiences, emotions, and meaning. On the other, NLP asks how some of that complexity can be translated into structures that machines can process.
That is what makes the fundamentals of NLP interesting to me.
NLP is not simply about making computers process words. The deeper challenge is how to bring structure, context, relationships, and meaning from human language into representations that machines can use.
And perhaps that was my main takeaway from the exam:
the more we try to make machines understand language, the more clearly we see how complex something as ordinary as human communication really is.
Join the conversation