Issue #151

AI Companion Devices Are Coming With Cameras Built In

Reports on OpenAI's new device and Meta's smart glasses raise questions about what always-on companions actually collect.

AI & TechAI Companion Devices Are Coming With Cameras Built In

AI Companion Devices Are Coming With Cameras and Sensors Built In

Iron Man’s Jarvis and Samantha from the movie Her are both AI companions that understand a person’s situation and help them through the day. Looking at recent news on AI device development, I notice a similar ambition taking shape.

According to reports from July 2026, OpenAI is developing a screenless, speaker-shaped device. Cameras, sensors, moving mechanical parts, and email-based personalization features have all been mentioned. None of this has been confirmed as a finalized launch spec yet. Meta, meanwhile, is already selling smart glasses equipped with a camera.

Reading this news, I was reminded of the AI surveillance cameras in a wealthy Toronto neighborhood that I covered a few months back. Back then, the issue was who could access the data from cameras installed for neighborhood safety. This time, what bothers me is that a device users carry on their body or keep in their home could be collecting information about their daily lives.

I think we need to scrutinize the scope of data personalization features require and how that data gets stored just as carefully as we scrutinize the product’s convenience.

inkling

Inkling: Our open-weights modelOur first open-weights model: multimodal, Mixture-of-Experts, with controllable reasoning effort. Available to fine-tune on Tinker.thinkingmachines.ai

The research directions are diverse, too. Mira Murati’s Thinking Machines released Inkling, an open-weights model that handles text, images, and voice, with adjustable settings for users. Yann LeCun, at AMI Labs, is focused on world models that predict how the world changes. These are different approaches, but they overlap in one shared interest: building AI that understands its surroundings. Neither approach necessarily requires a device that’s constantly filming the user.

Looking at this fragment, the translation is accurate, well-structured, and matches all counts (headings, footnotes, links, images, numbers). No Hangul remains, and all numbers are preserved as digits. No surgical edits are needed.

The reported companion device uses ambient context

OpenAI’s first hardware device is reportedly a screenless speaker that can move | TechCrunchA report on the development direction of a screenless companion device.techcrunch.com

The report describes a camera, environmental sensors, a portable battery, and a device that can move on its own. The idea is a product that becomes more personalized the longer it assists a user. OpenAI’s $6.5 billion acquisition of Jony Ive’s io team is itself a signal of the resources being poured into this new device. Still, we should read carefully — separating what’s confirmed, what’s planned, and what’s merely sourced from insiders.

Companion devices and surveillance devices can share components — cameras, microphones, sensors. But shared components don’t mean shared function or shared risk. What matters is when data is collected, whether it’s processed on-device, what gets sent to servers, and who has access.

Meta Ray-Ban Display — the innovative future of AI glasses - haebomA piece introducing the features of the Meta Ray-Ban Display and Neural Band.haebom.dev

Meta has been pushing in this direction much earlier, and much harder. Screenless smart glasses grew 167% by shipment volume in Q1 2026 alone — the fastest-growing wearable category — with Meta and EssilorLuxottica together holding more than 80% of the market. Last month, Meta launched a glasses model under its own “Meta” brand name rather than borrowing the Ray-Ban name, priced at $299 — $80 cheaper than the existing line. This can be read as a strategy to widen the user base through lower prices. Rising sales don’t mean every user’s footage gets collected as training data, so the actual data policy needs to be checked separately. The headline feature the glasses tout is “Super Sensing” — the ability to recognize objects and places in front of you in real time.

In the race for AI devices, how usefully a product grasps a user’s surroundings is becoming just as important as the quality of its answers. Smart glasses are one product form for capturing that context.

How World Model Research Connects to Camera Devices

You can see the same motivation—building the ability to understand one’s surroundings—in world model research too.

For Jarvis to actually be useful, being good at conversation isn’t enough. It needs to understand the layout of my room, where objects sit, and what’s happening right now. Yann LeCun believes today’s AI simply can’t do this. He recently stepped down as Meta’s Chief AI Scientist to found a company called AMI Labs (which announced roughly €890 million in funding in March 2026), and this was his opening argument: current models only appear intelligent because they handle language plausibly—they don’t actually understand the world.

The alternative LeCun is pursuing is world models1 and JEPA2—an approach aimed at building the ability to predict how objects change when they move or experience force. The underlying problem isn’t that physical laws are absent from text; it’s that language data alone has limits when it comes to learning the full variety of changes that occur in real environments. Fei-Fei Li of World Labs describes world models similarly, as an approach to learning the structure of space and time.

Video and sensor data can be used for this kind of learning. But since there are multiple possible data sources—public or licensed footage, simulations, and so on—world model research doesn’t necessarily mean constantly harvesting users’ daily lives. How personal data actually gets used in a real product is a separate matter of design and policy.

The two fields are related in that hardware captures data about one’s surroundings while the model interprets that data to help guide action. But that doesn’t mean we can conclude, just because one researcher changed jobs and started a company, that every AI device now needs to collect more surveillance data.

The English draft matches the Korean source accurately in structure, numbers, and meaning. No corrections needed.

We need to view both consumer demand and corporate platform competition together

There’s talk going around that “chatbots are over, and now everything’s moving off the screen,” but I think that’s an overstatement. You can see why by looking at what’s happening with immersive devices.

In the immersive device space, some development plans are being scaled back. Apple has scrapped development of a display for a budget XR headset it had been considering as a Vision Pro follow-up (an initiative the industry called G-VR), and Samsung Display also decided to shut the project down early in September. The $3,499 Vision Pro, meanwhile, has faced weak sales alongside reports of discontinuation and production halts. You can’t stretch the poor sales of one product into a claim that there’s no demand for immersive devices at all, but products with a high price tag and heavy wearing burden clearly haven’t found it easy to go mainstream.

Apple scraps development of ‘budget XR display’ - THE ELECA report covering the halt of Apple’s panel development project for a budget XR device.thelec.kr

Alongside headsets that take users into virtual space, companies are also developing glasses meant to be worn while looking at the real world. There’s a demand signal in the rising sales of smart glasses, but whether an “always-on companion” device will see long-term, widespread use still needs more evidence. There are cases like Humane’s AI Pin that were discontinued, and controversy has continued over the functionality and user experience of other devices.

I also think it’s important to note that companies are trying to stake out the dominant computing environment that comes after the smartphone. Meta’s wearables investment and OpenAI’s hardware development may well be driven by that kind of business calculation. Still, price tags or investment size alone don’t let us conclude that a company has given up on profit, or that a product exists purely to prop up an IPO. Factoring in entrants like Google and Samsung, what matters is whether a product’s convenience and its ability to retain customers actually translate into real results.

Meta announces new smart glasses starting at $299, as Zuckerberg keeps pushing wearablesMeta executives have said they see the lightweight smart glasses as a step towards a more advanced device that includes screens in the lenses.cnbc.com

I think the current competition is being driven by both consumers’ demand for convenience and companies’ strategies to lock down the next platform. It’s worth distinguishing how much the “companion” image companies project actually matches how their products are used in practice.

Game technology is also used to handle real-world environments

This kind of observation has actually been happening in games for a long time.

Games work by processing player input and changes in the in-game environment. Now, AI companions that react to gameplay situations are being layered on top of that. Krafton, working with Nvidia, has developed CPC (Co-Playable Character) technology and unveiled AI characters in “PUBG Ally” and inZOI. But playing a game doesn’t mean you’ve agreed to have every input collected and stored without limit.

Watch on YouTube

What’s interesting is what comes next. Krafton is now extending the physics-engine and virtual-environment capabilities it built for gaming into physical AI domains like robotics, autonomous driving, and defense. It has set up robotics research organizations in the US and Korea, teamed up with Hanwha Aerospace, and put ₩65,000,000,000 (~$47 million) into the car-sharing company SOCAR. Coverage of the deal has focused on the possibility of linking SOCAR’s roughly 1,100,000 kilometers of daily driving data to physical AI research. How much of that potential is actually realized will depend on the terms of the partnership and how the data is processed.

The simulation and environment-building capabilities accumulated through gaming could indeed be applied to training and testing robots. Still, we need to distinguish between the technology of building virtual environments and the act of collecting real-world life data. Using real-world data requires separately defining its purpose, scope, and access rights.

Oswarld’s Lens

When I’m mapping out a GTM strategy, I always ask two questions first: who is the real customer, and what is the real product? Apply these two questions to a companion device, and the answers get uncomfortable. On the surface, the customer looks like someone who wants a Jarvis-style AI. But what’s actually being sold might be the data generated by observing that user. This is exactly the structure Shoshana Zuboff described in The Age of Surveillance Capitalism — human experience becomes the raw material.

For personalization to work well, the system needs information about the user’s situation. The risk changes depending on whether that information stays on the device or gets transmitted to company servers. Unlike the neighborhood cameras discussed in the Rosedale piece, a device worn on the body can pick up information from much closer to daily life. I think that proximity demands equally easy visibility into, and control over, the scope of collection.

Devices like this could genuinely help people with mobility issues, seniors living alone, or anyone juggling multiple tasks at once. On-device3 processing is one way to cut down on external transmission. What worries me is a design where users have to hand over far more data than necessary just to use a convenience feature. Companies should have to explain exactly what information each feature requires, and users should be able to turn collection off or delete records.

So for me, the questions that matter as much as the launch itself are: when does the camera turn on, who stores and uses the data, and how can users actually control it?

Closing

The movie Her explores the intimacy an individual can feel toward AI. In real products, too, a friendly voice and responsive personality can build trust—but that alone doesn’t tell us whether our data is being handled safely.

When choosing a smart glasses or companion device, look at what tasks it helps with as well as what information it collects. Whether recording indicators, handling of bystanders’ information, server transmission and retention periods, and deletion features are clearly disclosed can serve as your criteria for judgment.

What feature would you find most useful in an AI companion device, and what information would you be reluctant to hand over?

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • AMI Labs, Official launch and funding announcement, 2026.3.10. ··· The funding raised was $1.03 billion, about €890 million.

  • Mark Gurman, “OpenAI’s First Device Will Be Moveable, Screenless Speaker Built as AI Companion”, Bloomberg, 2026. Related coverage ··· This is the original source for the spec claiming the device uses cameras, sensors, and email access, and even has its own motorized parts to move on its own. It’s the starting point for this piece’s premise that “companion = observation device.”

  • Kim Geon, “Apple Scraps ‘Entry-Level XR Display’ Development”, TheElec, 2026. Link ··· A report on Apple halting a specific XR display development project. It’s the evidence behind Part 3’s “from immersion to observation” argument.

  • “AI’s next frontier moves from words to world models”, Fortune, 2026. Link ··· A rundown of how LeCun and Fei-Fei Li each placed $1 billion bets on world models, along with their critiques of LLMs.

  • Fei-Fei Li & World Labs, “A Functional Taxonomy of World Models”, 2026. Link ··· A piece dividing world models into renderers, simulators, and planners. It explains the difference between the kind of information language models and world models each handle.

  • “Beyond Games — Krafton’s AI Territory Expands into Robots and Autonomous Driving”, News1, 2026. Link ··· The domestic-Korea evidence for the argument that game physics-engine capability is a “rehearsal” for physical AI.

Background / Concepts

  • Shoshana Zuboff, The Age of Surveillance Capitalism, 2019. ··· A book about how human experience becomes the raw material for prediction products. It’s the theoretical backbone of today’s piece.
  • “What is Joint Embedding Predictive Architecture (JEPA)?”, Turing Post, 2026. Link ··· A plain-language explainer, from I-JEPA to V-JEPA, of the concept that these models predict abstract representations rather than pixels.

Past pieces worth reading alongside this one

  • INLEVEL9 Letter, “How Much Is Your Privacy Worth” Read the past piece ··· Back then, surveillance still lived at the “neighborhood” scale. Think of this as the prequel to today’s piece — it lays the panopticon and surveillance-capitalism framework that today’s argument builds on.

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. World Model: An approach in which an AI builds an internal model of how the world works, so it can predict what happens next and plan actions accordingly. Think of it as learning the statistics of space and time, rather than the statistics of language.

  2. JEPA (Joint Embedding Predictive Architecture): An architecture trained to predict the “abstract representation” of missing parts within an observed context, rather than generating pixels or tokens one at a time. It’s the design LeCun champions as an alternative to LLMs.

  3. On-device: A method of processing data directly on the device itself rather than sending it to the cloud. Because the processed content never leaves the device, it’s relatively favorable from a privacy standpoint.