Voice Recognition and Touchless Technology

By | August 14, 2026
minipc2

Self-service technology is moving beyond the traditional touchscreen. In retail, hospitality, transportation, and food service, speech recognition and conversational AI and touchless interfaces are becoming important components of the next generation of kiosks.

The shift is not simply about removing physical contact. It is about creating faster, more flexible, and more accessible interactions between customers and automated systems.

From Touchscreens to Multimodal Interaction

For years, most self-service kiosks followed the same model: a display, touchscreen, payment system, and backend software.

That model remains effective, but AI is introducing new interaction methods.

Modern systems can combine:

  • Voice recognition
  • Touch interaction
  • Gesture control
  • Computer vision
  • QR and mobile interaction
  • AI-powered conversational interfaces

This creates a multimodal self-service system, allowing users to choose the interaction method that best fits the situation.

Recent deployments in Asia are already demonstrating voice-driven ordering and multilingual AI kiosk concepts, particularly in QSR and public-service environments.

What’s the difference between a demo-capable interface and a scalable self-service deployment.

A system that recognizes a coffee order in a quiet lab may fail in a food court because of music, fryer noise, queues, children, accents, multilingual code-switching, menu-specific terms, poor microphone placement, or two customers speaking nearby. The appropriate KPI is not simply speech-to-text accuracy. It is the end-to-end completion rate: correct order, correct modifiers, no staff intervention, acceptable elapsed time, no abandoned session, and no payment/order mismatch.

Accessibility and privacy

Voice is an additional channel, not an accessible replacement for touch. Same for gesture.

  • A voice-only interaction excludes people who cannot or prefer not to speak; audio-only output excludes deaf and hard-of-hearing users; and public voice interaction can be inappropriate for private tasks.

  • Accessible kiosks should provide independent alternatives across input and output: visual content, speech output, physical/tactile or other non-voice input, adjustable timing, and private listening where audio is used. The U.S. Access Board’s ICT requirements include private-listening capability for speech output and user-adjustable amplification/volume provisions in noisy public areas.access-board

  • W3C guidance similarly emphasizes that speech interaction needs well-labelled, controllable UI elements and alternative interaction paths.w3+1

  • Publicly visible microphones and cameras require transparent user notice: when they are active, what is captured, whether data is streamed or retained, who processes it, retention periods, and how users can use a non-voice/non-biometric alternative.

  • If voice is used as biometric identification rather than merely speech input, the article must treat that as a distinct, higher-risk design choice. Accessibility guidance for European requirements emphasizes protecting privacy when accessibility functions are used and offering alternatives to biometric identification/control.tetralogical+1

Voice Recognition for Self-Service

Voice recognition is particularly useful when customers need to complete complex tasks.

Instead of navigating through multiple menus, users can make requests using natural language.

For example:

“I want a large coffee with less ice and no sugar.”

The system can convert speech into text, identify the customer’s intent, process the requested options, and send the final order to the POS or backend system.

This is especially relevant for voice order systems in restaurants and drive-through environments.

The technology typically combines:

  • Microphone arrays
  • Automatic speech recognition
  • Natural language understanding
  • AI inference
  • Text-to-speech
  • POS integration

Noise cancellation and microphone positioning are particularly important in real-world deployments because restaurants, transport hubs, and retail stores are rarely quiet environments.

Core findings from TIG Report

  • Voice-AI penetration in drive-thrus should be counted by lane and microphone, not by outdoor-menu-board screen. Because a typical lane may have roughly three screens but one microphone, screen-based metrics can overstate penetration by up to 3×.

  • The report estimates around 80,000 U.S. drive-thru lanes have voice-capable infrastructure, largely associated with McDonald’s dual-lane footprint. This is hardware readiness—microphones, ODMBs, and headset systems—not evidence of AI order-taking at scale.

  • McDonald’s IBM automated-order pilot reportedly reached about 100 U.S. locations before being discontinued in 2024; the report characterizes its current AI-ordering layer as still under vendor evaluation.

  • New QSR kiosks are estimated to have only about 5% microphone capability, and retrofit adoption is below 1%. The report sees indoor kiosk voice primarily as an accessibility or narrowly guided-flow feature—not a broad replacement for touch.

  • The report puts cloud voice-AI costs at about $18,000–$20,000 per location annually for many enterprise deployments, though reported pricing ranges much more widely by vendor and deal structure.

  • A single-lane first-year deployment is estimated at:

    • $8,000–$25,000 for a clean retrofit.

    • $28,000–$65,000 for a new build.

    • $36,000–$93,000 for a complex retrofit.

  • The report disputes the common claim that voice AI simply “replaces a crew member.” In many drive-thrus, headset-enabled order taking is distributed across employees, so automation may lead to redeployment rather than an actual reduction in headcount.

  • Under the report’s example, an $18,000 annual SaaS fee at $0.75 of labor saved per order requires approximately 66 voice orders per day merely to cover SaaS, before hardware costs.

  • Retrofit and integration risk

    The report says that voice-AI deployments often become a cascading retrofit project:

    1. Legacy order-confirmation board becomes inadequate.

    2. ODMB hardware may need replacement.

    3. New Ethernet-capable conduit may be needed.

    4. Permitting and construction work may be triggered.

    5. Existing analog or legacy headset infrastructure may require modernization.

    6. Voice software is then added on top.

What Does Touchless Really Mean?

Touchless technology does not necessarily mean eliminating the display.

Instead, it can reduce or remove physical interaction with the screen through alternative interfaces.

Common approaches include:

Voice Interaction

Users speak directly to the kiosk instead of navigating menus manually.

Gesture Recognition

Cameras or sensors detect hand movements for simple selections.

Computer Vision

The system can detect users, products, or actions without requiring physical input.

Mobile Interaction

QR codes and smartphones can transfer part of the interaction from the kiosk to the customer’s own device.

This approach is already being used in smart reception, retail, and self-service environments.

No-Touch Touchscreens and Touchless Monitors

The concept of a no-touch touchscreen is particularly interesting for public environments.

Instead of requiring users to physically touch a display, cameras, infrared sensors, or other detection technologies can identify gestures or proximity.

For kiosk manufacturers, this creates new hardware requirements.

A touchless monitor may need:

  • High-brightness commercial display
  • Camera or depth sensor
  • Microphone array
  • Edge computing hardware
  • AI acceleration
  • Network connectivity

The display therefore becomes only one part of a larger AI interaction platform.

Applications in Retail and Hospitality

Retail and hospitality are among the strongest markets for these technologies.

Voice Order

Restaurants can use conversational AI to allow customers to place and customize orders through voice.

The system can also provide recommendations and upsell products during the conversation. Current commercial platforms are increasingly integrating voice ordering with POS, payment, and backend systems.

Hotel and Hospitality

Hotels can deploy voice-enabled kiosks for:

  • Check-in assistance
  • Wayfinding
  • Guest information
  • Service requests

This can reduce repetitive front-desk interactions while maintaining a digital customer service channel.

Retail

Retail kiosks can combine voice and computer vision to help customers locate products, check availability, and access personalized information.

Hardware Requirements Are Changing

The move toward AI-powered self-service also changes the hardware architecture.

A conventional kiosk may only require a CPU capable of running a web interface and digital signage software.

An AI-enabled kiosk may additionally require:

  • NPU or GPU acceleration
  • Multiple cameras
  • Microphone arrays
  • Local AI inference
  • Higher-performance embedded processors
  • Thermal management

Edge processing is particularly valuable when voice and vision workloads need low latency or when businesses want to minimize the amount of sensitive data sent to the cloud.

Why Asia Is an Important Market

Asia is well positioned for the development of touchless self-service because the region combines strong embedded hardware manufacturing with rapid adoption of retail automation and smart service technologies.

Manufacturers can integrate displays, embedded computers, cameras, sensors, microphones, and custom enclosures into a single OEM platform.

This makes Asia an important ecosystem for companies developing voice-enabled kiosks, smart retail terminals, and next-generation self-service hardware.

The Future of Touchless Self-Service

The future is unlikely to be completely touchscreen-free.

Instead, the industry is moving toward multimodal interaction.

A customer may speak to a kiosk, confirm the selection on the screen, scan a QR code with a phone, and complete payment through a contactless method.

The strongest systems will not necessarily replace one interface with another. They will combine different interfaces into a single customer journey.

Conclusion

Voice recognition, touchless technology, and AI are changing how people interact with self-service systems.

For kiosk manufacturers and system integrators, the opportunity is no longer simply to build a touchscreen terminal. The next generation of self-service hardware will combine displays, voice, computer vision, edge computing, and AI into a unified platform.

For Asian manufacturers in particular, this shift creates opportunities to develop more intelligent voice order, no-touch touchscreen, and touchless monitor solutions for retail, hospitality, transportation, and other commercial applications.

MORE ON TIG

Drive-thru versus kiosk

Dimension Drive-thru Dining-room kiosk
Customer expectation Speaking is already normal Touch is established behavior
Ambient environment Some vehicle isolation Music, crowds, kitchen noise
Privacy Relatively private lane Public shared setting
Vocabulary Menu-constrained Broader and less predictable
Retrofit economics Potentially viable Generally weak once touch works

The report sees kiosk voice as most defensible for:

  • ADA-oriented audio-guided ordering.

  • Constrained workflows such as pharmacy refill, check-in, or ticket redemption.

  • A touchless marketing position.

It expects the kiosk model through 2028 to be touch-first and voice-supported, not voice-replacing-touch.

What creates value

The report argues that the relevant KPI should be sales uplift, not labor savings. It cites kiosk deployments as delivering 15–30% sales uplift versus counter ordering, attributing the gain to a complete digital transaction experience:

  • Visible basket and order confirmation.

  • Well-timed upsells.

  • Menu and modifier logic.

  • Loyalty integration.

  • Payment integration.

  • Synchronized customer-facing display.

A voice solution that only automates the speaker post, without screen synchronization, basket visibility, loyalty, payment, and upsell logic, is framed as an incomplete product with limited ability to generate sales lift.

Vendor and market view

  • HME, PAR, and Panasonic form the microphone/headset layer.

  • Stratacache, Coates, Acrelec, and regional integrators support ODMB/display infrastructure.

  • Enterprise voice-AI vendors include SoundHound, Presto, Valyant AI, and Hi Auto.

  • Toast/Incept AI is positioned as an integrated drive-thru approach, but the report advises caution until field evidence validates headset integration, retrofit handling, and performance at varied site configurations.

  • For kiosk-focused deployments, it distinguishes:

    • Sodaclick: kiosk-native, signage-aware, edge/offline-capable.

    • SoundHound: enterprise conversational-AI platform with partial self-service integration.

    • ElevenLabs: high-quality speech synthesis API, but not an ordering platform; POS, menu, NLU, hardware, and commercial-flow integration must be built separately.

Case study: Taco Bell / Omilia

The report describes Taco Bell/Omilia as the largest U.S. production-scale stress test, citing 650+ Taco Bell locations in production as of a 2025 Omilia case study. Its assessment is that the deployment mainly replaces the speaker-post audio layer and lacks full screen, basket, and upsell integration. The report therefore classifies it as high deployment maturity but low integration depth, with no demonstrated sales lift.

Strategic outlook

The report’s 2026–2028 view:

  • Large-chain U.S. drive-thru voice: positive but measured growth.

  • International drive-thru voice: selective growth, constrained by local deployment realities.

  • Kiosk voice: flat, driven mostly by accessibility and limited guided flows.

  • SMB drive-thru voice: not broadly accessible yet due to cost and complexity.

  • Edge inference: the key structural opportunity because local inference could materially reduce cloud SaaS costs, though edge-GPU supply and hardware availability remain constraints.

VAIMS framework

The report introduces the Voice AI Implementation Maturity Score (VAIMS), a five-factor readiness assessment:

  1. Voice AI deployment maturity.

  2. Edge inference readiness.

  3. Integration depth and retrofit risk.

  4. Accuracy and fallback management.

  5. Cost-structure optimization.

Scores range from 0–25:

  • 0–10: High exposure.

  • 11–18: Transitional.

  • 19–25: Deployment-ready.

The report treats integration depth as the critical factor. It recommends verifying six production-ready elements:

  • POS/menu and modifier parity.

  • Loyalty and coupon support.

  • Payment integration.

  • CMS/menu synchronization.

  • Live order-confirmation screen synchronization.

  • Upsell logic plus documented ODMB/headset interoperability.

Bottom line

The report’s practical message is: do not buy “voice AI” as a standalone audio feature. Evaluate it as a complete, site-specific transaction-system upgrade with hard infrastructure dependencies, independently measured accuracy and escalation performance, and a credible route to lower-cost edge inference. For most operators, a 90-day, multi-location pilot with a control location is the appropriate proof threshold before scaling.