
Una interfaz de usuario de voz ( VUI ) permite la interacción verbal entre humanos y ordenadores, utilizando el reconocimiento de voz para comprender comandos hablados y responder preguntas , y normalmente la conversión de texto a voz para reproducir una respuesta. Un dispositivo de comandos de voz es un dispositivo controlado mediante una interfaz de usuario de voz.
Las interfaces de voz se han incorporado a automóviles , sistemas de domótica , sistemas operativos de computadoras , electrodomésticos como lavadoras y hornos microondas , y controles remotos de televisión . Son la principal forma de interactuar con asistentes virtuales en teléfonos inteligentes y altavoces inteligentes . Los sistemas de contestación automática más antiguos (que dirigen las llamadas a la extensión correcta) y los sistemas de respuesta de voz interactiva (que realizan transacciones más complejas por teléfono) pueden responder a la pulsación de botones del teclado mediante tonos DTMF , pero aquellos con una interfaz de voz completa permiten a quienes llaman expresar sus solicitudes y respuestas sin necesidad de pulsar ningún botón.
Los dispositivos de control por voz más recientes son independientes del hablante, por lo que pueden responder a múltiples voces, independientemente del acento o las influencias dialectales. También son capaces de responder a varios comandos a la vez, separando los mensajes de voz y proporcionando la retroalimentación adecuada , imitando con precisión una conversación natural. [ 1 ]
Descripción general
Una interfaz de voz (VUI) es la interfaz de cualquier aplicación de voz. Hasta hace poco, controlar una máquina simplemente hablándole solo era posible en la ciencia ficción . Hasta hace poco, este campo se consideraba inteligencia artificial . Sin embargo, los avances en tecnologías como la conversión de texto a voz, la conversión de voz a texto, el procesamiento del lenguaje natural y los servicios en la nube contribuyeron a la adopción masiva de este tipo de interfaces. Las VUI se han vuelto más comunes y las personas están aprovechando el valor que estas interfaces manos libres y sin necesidad de mirar ofrecen en muchas situaciones.
VUIs rely on the ability to process input reliably, inconsistent performance often leads to decreased user engagement and negative feedback. Designing a good VUI requires interdisciplinary talents of computer science, linguistics and human factors such as psychology. Even with advanced development tools, constructing an effective VUI requires understanding of both the tasks to be performed, as well as the target audience that will use the final system. The closer the VUI matches the user's mental model of the task, the easier it will be to use with little or no training, resulting in both higher efficiency and higher user satisfaction.
A VUI designed for the general public should emphasize ease of use and provide a lot of help and guidance for first-time callers. In contrast, a VUI designed for a small group of power users (including field service workers), should focus more on productivity and less on help and guidance. Such applications should streamline the call flows, minimize prompts, eliminate unnecessary iterations and allow elaborate "mixed initiative dialogs", which enable callers to enter several pieces of information in a single utterance and in any order or combination. In short, speech applications have to be carefully crafted for the specific business process that is being automated.
Not all business processes render themselves equally well for speech automation. In general, the more complex the inquiries and transactions are, the more challenging they will be to automate, and the more likely they will be to fail with the general public. In some scenarios, automation is simply not applicable, so live agent assistance is the only option. A legal advice hotline, for example, would be very difficult to automate. On the flip side, speech is perfect for handling quick and routine transactions, like changing the status of a work order, completing a time or expense entry, or transferring funds between accounts.
History
Early applications for VUI included voice-activated dialing of phones, either directly or through a (typically Bluetooth) headset or vehicle audio system.
In 2007, a CNN business article reported that voice command was over a billion dollar industry and that companies like Google and Apple were trying to create speech recognition features.[2] In the years since the article was published, the world has witnessed a variety of voice command devices. Additionally, Google has created a speech recognition engine called Pico TTS and Apple released Siri. Voice command devices are becoming more widely available, and innovative ways for using the human voice are always being created. For example, Business Week suggests that the future remote controller is going to be the human voice. Currently Xbox Live allows such features and Jobs hinted at such a feature on the new Apple TV.[3]
Voice command software products on computing devices
Both Apple Mac and WindowsPC provide built in speech recognition features for their latest operating systems.
Microsoft Windows
Two Microsoft operating systems, Windows 7 and Windows Vista, provide speech recognition capabilities. Microsoft integrated voice commands into their operating systems to provide a mechanism for people who want to limit their use of the mouse and keyboard, but still want to maintain or increase their overall productivity.[4]
Windows Vista
With Windows Vista voice control, a user may dictate documents and emails in mainstream applications, start and switch between applications, control the operating system, format documents, save documents, edit files, efficiently correct errors, and fill out forms on the Web. The speech recognition software learns automatically every time a user uses it, and speech recognition is available in English (U.S.), English (U.K.), German (Germany), French (France), Spanish (Spain), Japanese, Chinese (Traditional), and Chinese (Simplified). In addition, the software comes with an interactive tutorial, which can be used to train both the user and the speech recognition engine.[5]
Windows 7
In addition to all the features provided in Windows Vista, Windows 7 provides a wizard for setting up the microphone and a tutorial on how to use the feature.[6]
Mac OS X
Todos los ordenadores Mac OS X vienen con el software de reconocimiento de voz preinstalado. Este software es independiente del usuario y permite navegar por menús e introducir atajos de teclado; leer nombres de casillas de verificación, botones de opción, elementos de listas y botones; y abrir, cerrar, controlar y alternar entre aplicaciones. [ 7 ] Sin embargo, el sitio web de Apple recomienda adquirir un producto comercial llamado Dictate . [ 7 ]
Productos comerciales
Si un usuario no está satisfecho con el software de reconocimiento de voz integrado o no dispone de un software de reconocimiento de voz integrado en su sistema operativo, puede experimentar con un producto comercial como Braina Pro o DragonNaturallySpeaking para PC con Windows, [ 8 ] y Dictate, el mismo software para Mac OS. [ 9 ]
Dispositivos móviles controlados por voz
Cualquier dispositivo móvil con sistema operativo Android, Microsoft Windows Phone, iOS 9 o posterior, o Blackberry OS ofrece funciones de comandos de voz. Además del software de reconocimiento de voz integrado en el sistema operativo de cada teléfono móvil, el usuario puede descargar aplicaciones de comandos de voz de terceros desde la tienda de aplicaciones de cada sistema operativo: Apple App Store , Google Play , Windows Phone Marketplace (inicialmente Windows Marketplace for Mobile ) o BlackBerry App World .
Sistema operativo Android
Google ha desarrollado un sistema operativo de código abierto llamado Android , que permite al usuario realizar comandos de voz como: enviar mensajes de texto, escuchar música, obtener direcciones, llamar a empresas, llamar a contactos, enviar correos electrónicos, ver un mapa, ir a sitios web, escribir una nota y buscar en Google. [ 10 ] El software de reconocimiento de voz está disponible para todos los dispositivos desde Android 2.2 "Froyo" , pero la configuración debe estar en inglés. [ 10 ] Google permite al usuario cambiar el idioma, y se le pregunta al usuario la primera vez que usa la función de reconocimiento de voz si desea que sus datos de voz se asocien a su cuenta de Google. Si un usuario decide optar por este servicio, permite a Google entrenar el software con la voz del usuario. [ 11 ]
Google presentó el Asistente de Google con Android 7.0 "Nougat" . Es mucho más avanzado que la versión anterior.
Amazon.com ofrece el Echo , que utiliza la versión personalizada de Android de Amazon para proporcionar una interfaz de voz.
Microsoft Windows
Windows Phone is Microsoft's mobile device's operating system. On Windows Phone 7.5, the speech app is user independent and can be used to: call someone from your contact list, call any phone number, redial the last number, send a text message, call your voice mail, open an application, read appointments, query phone status, and search the web.[12][13] In addition, speech can also be used during a phone call, and the following actions are possible during a phone call: press a number, turn the speaker phone on, or call someone, which puts the current call on hold.[13]
Windows 10 introduces Cortana, a voice control system that replaces the formerly used voice control on Windows phones.
iOS
Apple added Voice Control to its family of iOS devices as a new feature of iPhone OS 3. The iPhone 4S, iPad 3, iPad Mini 1G, iPad Air, iPad Pro 1G, iPod Touch 5G and later, all come with a more advanced voice assistant called Siri. Voice Control can still be enabled through the Settings menu of newer devices. Siri is a user independent built-in speech recognition feature that allows a user to issue voice commands. With the assistance of Siri a user may issue commands like, send a text message, check the weather, set a reminder, find information, schedule meetings, send an email, find a contact, set an alarm, get directions, track your stocks, set a timer, and ask for examples of sample voice command queries.[14] In addition, Siri works with Bluetooth and wired headphones.[15]
Apple introduced Personal Voice as an accessibility feature in iOS 17, launched on September 18, 2023.[16] This feature allows users to create a personalized, machine learning-generated (AI) version of their voice for use in text-to-speech applications. Designed particularly for individuals with speech impairments, Personal Voice helps preserve the unique sound of a user's voice. It enhances Siri and other accessibility tools by providing a more personalized and inclusive user experience. Personal Voice reflects Apple's ongoing commitment to accessibility and innovation.[17][18]
Amazon Alexa
In 2014 Amazon introduced the Alexa smart home device. Its main purpose was just a smart speaker, that allowed the consumer to control the device with their voice. Eventually, it turned into a novelty device that had the ability to control home appliance with voice. Now almost all the appliances are controllable with Alexa, including light bulbs and temperature. By allowing voice control, Alexa can connect to smart home technology allowing users to lock their house, control the temperature, and activate various devices. This form of A.I allows for someone to simply ask it a question, and in response Alexa searches for, finds, and recites the answer back to them.[19]
Speech recognition in cars
As car technology improves, more features will be added to cars and these features could potentially distract a driver. Voice commands for cars, according to CNET, should allow a driver to issue commands and not be distracted. CNET stated that Nuance was suggesting that in the future they would create a software that resembled Siri, but for cars.[20] Most speech recognition software on the market in 2011 had only about 50 to 60 voice commands, but Ford Sync had 10,000.[20] However, CNET suggested that even 10,000 voice commands was not sufficient given the complexity and the variety of tasks a user may want to do while driving.[20] Voice command for cars is different from voice command for mobile phones and for computers because a driver may use the feature to look for nearby restaurants, look for gas, driving directions, road conditions, and the location of the nearest hotel.[20] Currently, technology allows a driver to issue voice commands on both a portable GPS like a Garmin and a car manufacturer navigation system.[21]
List of Voice Command Systems Provided By Motor Manufacturers:
- Ford Sync
- Lexus Voice Command
- Chrysler UConnect
- Honda Accord
- GM IntelliLink
- BMW
- Mercedes
- Pioneer
- Harman
- Hyundai
Non-verbal input
While most voice user interfaces are designed to support interaction through spoken human language, there have also been recent explorations in designing interfaces take non-verbal human sounds as input.[22][23] In these systems, the user controls the interface by emitting non-speech sounds such as humming, whistling, or blowing into a microphone.[24]
One such example of a non-verbal voice user interface is Blendie,[25][26] an interactive art installation created by Kelly Dobson. The piece comprised a classic 1950s-era blender which was retrofitted to respond to microphone input. To control the blender, the user must mimic the whirring mechanical sounds that a blender typically makes: the blender will spin slowly in response to a user's low-pitched growl, and increase in speed as the user makes higher-pitched vocal sounds.
Another example is VoiceDraw,[27] a research system that enables digital drawing for individuals with limited motor abilities. VoiceDraw allows users to "paint" strokes on a digital canvas by modulating vowel sounds, which are mapped to brush directions. Modulating other paralinguistic features (e.g. the loudness of their voice) allows the user to control different features of the drawing, such as the thickness of the brush stroke.
Other approaches include adopting non-verbal sounds to augment touch-based interfaces (e.g. on a mobile phone) to support new types of gestures that wouldn't be possible with finger input alone.[24]
Design challenges
Voice interfaces pose a substantial number of challenges for usability. In contrast to graphical user interfaces (GUIs), best practices for voice interface design are still emergent.[28]
Discoverability
With purely audio-based interaction, voice user interfaces tend to suffer from low discoverability:[28] it is difficult for users to understand the scope of a system's capabilities. In order for the system to convey what is possible without a visual display, it would need to enumerate the available options, which can become tedious or infeasible. Low discoverability often results in users reporting confusion over what they are "allowed" to say, or a mismatch in expectations about the breadth of a system's understanding.[29][30]
Transcription
While speech recognition technology has improved considerably in recent years, voice user interfaces still suffer from parsing or transcription errors in which a user's speech is not interpreted correctly.[31] These errors tend to be especially prevalent when the speech content uses technical vocabulary (e.g. medical terminology) or unconventional spellings such as musical artist or song names.[32]
Understanding
El diseño eficaz de sistemas para maximizar la comprensión conversacional sigue siendo un área de investigación abierta. Las interfaces de usuario de voz que interpretan y gestionan el estado conversacional son difíciles de diseñar debido a la dificultad inherente de integrar tareas complejas de procesamiento del lenguaje natural, como la resolución de correferencias , el reconocimiento de entidades nombradas , la recuperación de información y la gestión del diálogo . [ 33 ] La mayoría de los asistentes de voz actuales son capaces de ejecutar comandos individuales muy bien, pero su capacidad para gestionar el diálogo más allá de una tarea específica o un par de turnos en una conversación es limitada. [ 34 ]
Implicaciones para la privacidad
La privacidad se ve afectada por el hecho de que los proveedores de interfaces de usuario de voz tienen acceso a los comandos de voz sin cifrar, pudiendo así compartirlos con terceros y procesarlos de forma no autorizada o inesperada. [ 35 ] [ 36 ] Además del contenido lingüístico del habla grabada, la forma de expresión y las características de la voz del usuario pueden contener implícitamente información sobre su identidad biométrica, rasgos de personalidad, complexión, estado de salud física y mental, sexo, género, estados de ánimo y emociones , estatus socioeconómico y origen geográfico. [ 37 ]
Véase también
- Síntesis de voz
- Lista de software de reconocimiento de voz
- Interfaz de usuario en lenguaje natural
- Diseño de interfaz de usuario
- Navegador de voz
- Reconocimiento de voz en Linux
- Linguatronic
- Computación por voz
Referencias
- ^ "Control por voz de la lavadora" . Revista Appliance Magazine . Archivado del original el 3 de noviembre de 2011. Consultado el 20 de diciembre de 2018 .
- ^ Borzo, Jeanette (8 de febrero de 2007). "Ahora sí que hablas" . CNN Money. Archivado del original el 11 de febrero de 2007. Recuperado el 25 de abril de 2012 .
- ^ "¿Control por voz, el fin del control remoto del televisor?" . Bloomberg.com . Business Week. 9 de diciembre de 2011. Archivado del original el 8 de diciembre de 2011. Consultado el 1 de mayo de 2012 .
- ^ "Windows Vista Built In Speech" . Windows Vista . Consultado el 25 de abril de 2012 .
- ^ "Funcionamiento por voz en Vista" . Microsoft.
- ^ "Configuración del reconocimiento de voz" . Microsoft.
- ^ a b "Habilidades físicas y motoras" . Apple.
- ^"DragonNaturallySpeaking PC". Nuance.
- ^"DragonNaturallySpeaking Mac". Nuance.
- ^ ab"Voice Actions".
- ^"Google Voice Search For Android Can Now Be "Trained" To Your Voice". 14 December 2010. Retrieved 24 April 2012.
- ^"Using Voice Command". Microsoft. Retrieved 24 April 2012.
- ^ ab"Using Voice Commands". Microsoft. Retrieved 27 April 2012.
- ^"Siri, The iPhone 3GS & 4, iPod 3 & 4, have voice control like an express Siri, it plays music, pauses music, suffle, Facetime, and calling Features". Apple. Retrieved 27 April 2012.
- ^"Siri FAQ". Apple.
- ^"How to use Personal Voice on iPhone with iOS 17". Engadget. 2023-12-06. Retrieved 2024-08-21.
- ^Jason England (2023-07-13). "How to set up and use Personal Voice in iOS 17 — make your iPhone sound just like you". LaptopMag. Retrieved 2024-08-21.
- ^"Advancing Speech Accessibility with Personal Voice". Apple Machine Learning Research. Retrieved 2024-08-21.
- ^"How Amazon's Echo went from a smart speaker to the center of your home". Business Insider.
- ^ abcd"Siri Like Voice". CNET.
- ^"Portable GPS With Voice". CNET.
- ^Blattner, Meera M.; Greenberg, Robert M. (1992). "Communicating and Learning Through Non-speech Audio". Multimedia Interface Design in Education. pp. 133–143. doi:10.1007/978-3-642-58126-7_9. ISBN 978-3-540-55046-4.
- ^Hereford, James; Winn, William (October 1994). "Non-Speech Sound in Human-Computer Interaction: A Review and Design Guidelines". Journal of Educational Computing Research. 11 (3): 211–233. doi:10.2190/mkd9-w05t-yj9y-81nm. ISSN 0735-6331. S2CID 61510202.
- ^ abSakamoto, Daisuke; Komatsu, Takanori; Igarashi, Takeo (27 August 2013). "Voice augmented manipulation | Proceedings of the 15th international conference on Human-computer interaction with mobile devices and services": 69–78. doi:10.1145/2493190.2493244. S2CID 6251400. Retrieved 2019-02-27.
{{cite journal}}: Cite journal requires|journal=(help) - ^Dobson, Kelly (August 2004). "Blendie | Proceedings of the 5th conference on Designing interactive systems: processes, practices, methods, and techniques": 309. doi:10.1145/1013115.1013159. Retrieved 2019-02-27.
{{cite journal}}: Cite journal requires|journal=(help) - ^"Kelly Dobson: Blendie". web.media.mit.edu. Archived from the original on 2022-05-10. Retrieved 2019-02-27.
- ^Harada, Susumu; Wobbrock, Jacob O.; Landay, James A. (15 October 2007). "Voicedraw | Proceedings of the 9th international ACM SIGACCESS conference on Computers and accessibility": 27–34. doi:10.1145/1296843.1296850. S2CID 218338. Retrieved 2019-02-27.
{{cite journal}}: Cite journal requires|journal=(help) - ^ abMurad, Christine; Munteanu, Cosmin; Clark, Leigh; Cowan, Benjamin R. (3 September 2018). "Design guidelines for hands-free speech interaction | Proceedings of the 20th International Conference on Human-Computer Interaction with Mobile Devices and Services Adjunct": 269–276. doi:10.1145/3236112.3236149. S2CID 52099112. Retrieved 2019-02-27.
{{cite journal}}: Cite journal requires|journal=(help) - ^Yankelovich, Nicole; Levow, Gina-Anne; Marx, Matt (May 1995). "Designing SpeechActs | Proceedings of the SIGCHI Conference on Human Factors in Computing Systems": 369–376. doi:10.1145/223904.223952. S2CID 9313029.
{{cite journal}}: Cite journal requires|journal=(help) - ^"What can I say? | Proceedings of the 18th International Conference on Human-Computer Interaction with Mobile Devices and Services". doi:10.1145/2935334.2935386. S2CID 6246618.
{{cite journal}}: Cite journal requires|journal=(help) - ^Myers, Chelsea; Furqan, Anushay; Nebolsky, Jessica; Caro, Karina; Zhu, Jichen (19 April 2018). "Patterns for How Users Overcome Obstacles in Voice User Interfaces | Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems": 1–7. doi:10.1145/3173574.3173580. S2CID 5041672. Retrieved 2019-02-27.
{{cite journal}}: Cite journal requires|journal=(help) - ^Springer, Aaron; Cramer, Henriette (21 April 2018). ""Play PRBLMS" | Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems": 1–13. doi:10.1145/3173574.3173870. S2CID 5050837. Retrieved 2019-02-27.
{{cite journal}}: Cite journal requires|journal=(help) - ^Galitsky, Boris (2019). Developing Enterprise Chatbots: Learning Linguistic Structures (1st ed.). Cham, Switzerland: Springer. pp. 13–24. doi:10.1007/978-3-030-04299-8. ISBN 978-3-030-04298-1. S2CID 102486666.
- ^Pearl, Cathy (2016-12-06). Designing Voice User Interfaces: Principles of Conversational Experiences (1st ed.). Sebastopol, CA: O'Reilly Media. pp. 16–19. ISBN 978-1-491-95541-3.
- ^"Apple, Google, and Amazon May Have Violated Your Privacy by Reviewing Digital Assistant Commands". Fortune. 2019-08-05. Retrieved 2020-05-13.
- ^Hern, Alex (2019-04-11). "Amazon staff listen to customers' Alexa recordings, report says". the Guardian. Retrieved 2020-05-21.
- ^Kröger, Jacob Leon; Lutz, Otto Hans-Martin; Raschke, Philip (2020). "Privacy Implications of Voice and Speech Analysis – Information Disclosure by Inference". Privacy and Identity Management. Data for Better Living: AI and Privacy. IFIP Advances in Information and Communication Technology. Vol. 576. pp. 242–258. doi:10.1007/978-3-030-42504-3_16. ISBN 978-3-030-42503-6. ISSN 1868-4238.
External links
- Voice Interfaces: Assessing the Potential by Jakob Nielsen
- The Rise of Voice: A Timeline
- Voice First Glossary of Terms
- Voice First A Reading List
- User interface techniques
- Voice technology
- Speech recognition
- History of human–computer interaction
- Human–computer interaction
- Computing input devices
- Assistive technology
- Speech processing
- Ambient intelligence