A featured contribution from Leadership Perspectives, a curated forum for startup ecosystem leaders, nominated by our subscribers and vetted by the Startup City Editorial Board.

SeedIL Ventures

The Invisible Part of AI

Cynthia Phitoussi

While it is very common to view AI agents as intelligent machines, the reality is pretty far from that as AI lacks a critical component that every human possesses: reasoning.

What is reasoning?

Let’s take the following simple sentence: “I threw the ball to the sky and it fell”.

Anyone of us, even a child, will immediately understand the meaning of this sentence. However, an AI agent can also easily reach the following conclusions:

• The sky can fall
• The ball is now staying up in the air.
• I managed to make the sky fall using a ball.

Why do we, humans, simply “get the right context”?

We have basic knowledge about the way the world works. We know gravity makes things go up and down, and we know that even though “fallen skies” exists in our language as a metaphor, skies can’t physically fall. This is very basic reasoning, yet even this basic reasoning is above existing AI capabilities.

AI Agents are remarkable copiers, yet they don’t understand the reasoning behind the intelligence they copy.

So, how does AI really work?

AI products are essentially intelligence copy machines: every AI-powered product contains one or several “models”, which are essentially pieces of machine-generated code that are trained based on data in a process called data labelling/model training.

Indeed, today’s smart machines don’t have actual intelligence. Instead they behave like the perfect copycat. AI products are usually designed to answer questions in some way: from medical diagnosis to self-driving cars, all AI models are question/answer machines. In the AI learning process, many examples (typically millions) are gathered and the machine is trained on them, memorizing the answers for a given question. If we take as an example a company wishing to develop a skin cancer detection mobile app, they will need many images of moles to be collected with a Positive/Negative label. Following the training processes, the AI agent (mobile app in this case) will be able to predict the risk of a new mole image based on the past data it was exposed to. The above process contains 3 components:

1. Data collection - The skin cancer app developers need access to millions of images of moles.

2. Data labelling - For each image, a human data labeler, a doctor in this case, should indicate whether the mole is malignant suspected, or benign.

3. Model training - once the data has been labeled, it is used to produce a model that outputs malignant or benign based on a new image.

As the above process demonstrates, the skin cancer detector has no real idea of what a mole is, it is merely copying the intelligence of human doctors that was embedded into the data in the form of labels. Therefore, if we want to develop face detection, car detection or product detection, the exact process will take place only with different data and different labels.

A new industry is born

While most of the conversation around AI is about the researchers who are developing the algorithms, AI product development cost, time and accuracy is centered around the data, taking up more than 85% of the R&D.

With the growing market demand for smart products, a new industry is born: the data labeling industry. Expected to reach $8B by 2025, this industry is already at the core of leading AI companies and the numbers are pretty staggering. AI products like Alexa, Tesla, Google and many others are each employing thousands of people who label data all day long and essentially play the role of machines’ teachers.

A16Z have conducted a great analysis of this industry, stating that AI companies appear increasingly, to combine elements of both software and services with gross margins, scaling, and defensibility that may represent a new class of business entirely.

At SeedIL Ventures, we entered the field of AI labelling a year ago when we invested in the fast growing Israeli startup, Dataloop. Dataloop’s platform covers the entire data preparation cycle, from data labeling, automating data ops, to customizing production pipelines, and back to weaving the human-in-the-loop. Today, Dataloop addresses all kinds of clients and industries from video moderation, self-checkout and smart scales solution in retail, to security and autonomous cars.

So, is it all just hype?

It is clear that AI is over hyped, yet the answer divides into two parts.

Can AI reach human levels of intelligence in the near future? No, it will likely take decades and still will be missing some key components that we have not managed to figure out yet.

Will AI disrupt industries in the near future? Yes, even the process of intelligence copy is a major industrial breakthrough, opening the doors to massive automation across verticals: autonomous driving, automatic medical diagnosis, industrial inspection and self-checkout retail to name a few. Data and domain expertise are becoming the fuel for machine automation, businesses which fail to adopt this new paradigm will be unable to compete within few years.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.

Weekly Brief