Powered by Blogger.
Showing posts with label Acoustic processing. Show all posts
Showing posts with label Acoustic processing. Show all posts

The 10 qualities of highly effective hands-free systems

Monday, January 28, 2013

The first time I saw — and heard — a hands-free kit in action was in 1988. (Or was it 1989? Meh, same difference.) At the time, I was pretty impressed with the sound quality. Heck, I was impressed that hands-free conversations were even possible. You have to remember that mobile phones were still an expensive novelty — about $4000 in today’s US dollars. And good grief, they looked like this:



It’s almost a shock to see how far we’ve come since 1988. We’ve become conditioned to devices that cost far less, do far more, and fit into much smaller pockets. (Though, admittedly, the size trend for smartphones has shifted into reverse.) Likewise, we’ve become conditioned to hands-free systems whose sound quality would put that 1998 kit to shame. The sound might have been okay at the time, but because of the contrast effect, it wouldn’t pass muster today. Our ears have become too discerning.

Which brings me to a new white paper from Phil Hetherington and Andrew Mohan of the acoustics team at QNX Software Systems. Evaluating hands-free solutions from various suppliers can be a complex endeavor, for the simple fact that hands-free systems have become so sophisticated and complex. To help simplify the decision process, Phil and Andrew have boiled the problem down to 10 key factors:

  • Acoustic echo cancellation
  • Noise reduction and speech reconstruction
  • Multi-channel support
  • Automatic gain control
  • Equalization
  • Wind buffet suppression
  • Intelligibility enhancement
  • Noise dependent receive gain
  • Bandwidth extension
  • Wideband support

Ultimately, you must judge a hands-free solution by the quality of the useful sound it delivers. By focusing on these 10 essentials, you can make a much sounder judgment (pun fully intended).

Recently, Electronic Design published a version of this paper on their website. For a longer version, which includes a decision checklist, visit the QNX download center.

Full-duplex in the car? Who even knows what this means?

Thursday, January 10, 2013

I know lots of people don't understand full-duplex. Hell, I think most people have never even heard of it. Unfortunately, these same people have repeatedly experienced its poor cousin — half-duplex — without really understanding either.

Please don't take me wrong. I'm not patronizing. I spent a few years in telecoms and barely understand it myself. What I do know is that half-duplex = bad. And that full-duplex = good.

So when I talked to my colleagues at the office today, I knew we were using half-duplex. How did I know? I started to say something and so did someone else at the other end of the line. I couldn't hear them talking while I was talking (a certain amount of latency adds to the circus) so I stopped talking... and so did they. Then there were lots of simultaneous barely-understood apologies and a long uncomfortable silence. Then we both tried to break the silence at the same time; more uncomfortable silence. Very awkward and distracting. I know you know what I mean.

So... um... could someone please fix this? I mean, we send people to the moon after all.

I hate to blow our own (proverbial) horn (well sometimes) but believe me, I have to. QNX has THE best audio technology solution in or out of the car. And this technology, just so happens to be in the new QNX technology concept car at CES. Yes, really!

Today I witnessed a conversation in the QNX booth between someone in the Bentley and someone in a sound-proof booth. Well, to my (sheer) delight, one person talked while the other person talked over him... and both heard the other! Just like a real face-to-face conversation with overlapping dialogue. It was so natural, it almost slipped by as if it were expected:



In a world where communication is more often than not at the root of all successes and failures, I think this is nothing short of a long-overdue miracle.

Am I crazy for talking to my car?

Wednesday, August 15, 2012

Earlier this afternoon, I participated in a connected car panel at SpeechTEK 2012, hosted by our friend Mazin Gilbert from AT&T. The other panelists included Greg Bielby of VoltDelta, Thomas Schalk of Agero, and Hakan Kostepen of Panasonic.

Even though Mazin did a fantastic job, not every panelist had a chance to answer every question. I was itching to answer some, so here are my responses to the questions that I didn't get to answer, or where I feel I could have provided a more complete response.

Have speech technologies matured to the point where they can be used robustly in the car? The general answer to this question from the panel was yes, but I think the real answer is a qualified yes. The technologies exist, but often aren't applied or may need auto-specific adaptations to handle in-cabin noise or other issues. Natural language recognition was an oft-stated driving technology, but a missing piece to the puzzle is hybrid recognition. I don't mean pushing recognition wholesale to the cloud, like Siri does. I mean a true split of the recognition effort, where each part does what it’s best at. Put the front half of acoustic processing in the vehicle to clean up the audio and convert the waveform to frequency-domain data, then send the data to the cloud-based server. The cloud server can then parse and interpret the data, and send back the result.

Hybrid speech rec solves three problems at once: better audio signals (the car can improve audio specific to the in-cabin environment), better cost (frequency data is far more compressed than raw audio, so you pay less for data transfer), and better responsiveness (hybrid rec gives the server time to start working on the response while it's coming in instead of waiting for the whole utterance to finish before starting).

Is driver distraction a major business driver, or is it the "Siri effect"? Currently, the car industry seems to use driver distraction as a reason to push a lot of features into speech. Many of those uses are gimmicky. Personally, I don't care if I can set my climate control system with voice — why would I when I can simply turn a dial? I once had someone ask me about the feasibility of adding voice recognition commands for rolling down the windows. I asked him, "Yes, but wouldn't people just push the window button?"

We shouldn’t implement speech commands just because we can. They may have contributed to excitement in the early adopter crowd, but we're beyond that now. Mind you, there are some seriously useful ways to use voice. For instance, any time you need to pick from a huge number of choices, voice recognition is the natural way to go. Calling contacts ("Call Sarah Potter"), entering destinations ("Go to 3121 South Park Street"), or picking music ("Play Audioslave") are all much easier than using an HMI to enter the same information, and safer to boot. It just has to work consistently and accurately.

Will car makers see more speech moving to the cloud, or will it be a hybrid of cloud and embedded? I disagree with the majority of the panel on this one, and, I think, the majority of people in the industry. Most auto people believe a hybrid between embedded and cloud allows the best of both worlds — good recognition and updatability when connected, and consistent reliability when not. My colleague Andrew Poliak also champions this view with a memorable catch phrase: Zombie Apocalypse. That is, you still want the system to work, albeit partially, when the infrastructure isn't available.

But if you ask me, everyone is missing the point — theirs is a technology-centric point of view. Everyday customer acceptance of a particular technology is notoriously harsh: if it doesn't work well, it gets rejected out of hand. Good cloud solutions beat an embedded solution hands-down; they just need some improvements (see my hybrid bullet above). Once a customer experiences a good solution, they will become frustrated with one that performs poorly. In my opinion, it's better not to offer the service at all, than to try a graceful degradation of capability, because most customers won't understand or care. Spend the effort instead on making sure you always have an acceptable cloud connection — either through multiple redundant mechanisms or a car-based powerful antenna — and you'll be better off. Even when the car knows some data that the cloud doesn't (like a mobile's contact list or music selection), there's no need to handle that on the embedded side. The cloud recognition server is powerful enough to not require the data set a priori. And I think we can predict an eventual migration of phone data to cloud-based data (or cloud-synchronized data) that makes the car's knowledge either easily transferrable or less relevant.

Who makes money, and how, from voice-enabled agents or voice services? This was one of the best questions of the panel, because nobody really knows the exact model, but everybody agreed that customer tolerance is very low. The most likely candidate is ad-based revenue. This doesn't mean reading ads aloud to the driver, but rather, positively influencing search results for either active or temporary situation-based points of interest (POIs). Depending on how valuable the service is to the driver, there will still be an option for service-based payments and high-value apps.

Standards and building mobile apps — will it come? You need standards if you want to build an app platform that will promote application creation and adoption. That's what we're doing with the QNX CAR 2 application platform — creating a way for someone other than the car companies to join the ecosystem and to deploy their apps to the car in a controlled way. But don't forget, you need a standard way to deploy apps for the cloud half of the recognition, too.

To close, let me share two photos. One was taken outside the Marriott Marquis, the hotel hosting the conference just off of Times Square in NYC. The other is from our PR agency, Breakaway Communications. What do they have in common? Wooden water towers. Sorry, I couldn't help myself; I just love those things. They just look so quaint in a city full of glass and brick.






In-car displays you hear, rather than see

Tuesday, June 19, 2012

We still have a lot in common with our caveman ancestors. (Yes, I know, they didn't all live in caves. Some lived in forests, others in savannahs, and still others in jungles. But I'm trying to make a point, so bear with me!)

Take, for example, our sense of hearing. At one time, we used auditory cues to locate prey or, conversely, avoid becoming prey. If a cave bear growled, getting a fix on the location of the growl could mean the difference between life and death. At the very least, it helped you avoid running directly into the bear's mouth.

Kidding aside, the human auditory system has a serious ability to fix the location, direction, and trajectory of objects, be they cave bears or Buicks. And it's an ability that's been honed from time immemorial. So why not take advantage of it when creating user interfaces for cars?

Which brings us to spatial auditory displays. In a nutshell, these displays allow you to perceive sound as coming from various locations in a three-dimensional space. Deployed in a car, they can help you intuitively identify voices and sources of instructions, and help pinpoint the location and relative trajectory of danger. They can also improve reaction times to application prompts and potentially hazardous events.
Interested in this topic? Learn more in Scott Pennock's ECD article, "Spatial auditory displays: Reducing cognitive load and improving driver reaction times."

I know, that's a lot to take in. So let's look at an example.

Locating the emergency vehicle, without really trying
Have you ever been cruising along when, suddenly, you hear an ambulance siren? I don't know about you, but I often spend time figuring out where, exactly, the ambulance is coming from. And I don't always get it right. That's called a location error.

Such errors can occur for a variety of reasons. For example, if the ambulance is approaching from the right, but your left window is open and a building on the left is reflecting sound from the siren, you might make the mistake of thinking that the ambulance is approaching from the left. Your mind realizes, quite correctly, that the sound is coming from the left, but the environment is conspiring to mask where the sound is actually coming from.

A spatial auditory display can help address this problem by controlling the acoustic cues you hear. The degree to which the display can do this depends, in part, on the hardware employed. For example, a display based on a large array of loudspeakers can provide more location information than one based on two loudspeakers.

In any case (and this is important), the display can help you determine the location more quickly and with less cognitive load — which means you may have more brain cycles to respond to the situation appropriately.


Helping the driver locate and track an emergency vehicle

A slight right, not a sharp right
I'm only scratching the surface here. Spatial auditory displays can, in fact, help improve all kinds of driving activities, from engaging in a handsfree call to using your navigation system.

For example, rather than simply say "turn right", the display could emit the instruction from the right side of the vehicle. It could even use apparent motion of the auditory prompt to convey a slight right as opposed to a sharp right.

But enough from me. To learn more about spatial auditory displays, check out a new article from my colleague Scott Pennock, whose knowledge of spatial auditory displays far surpasses mine. The article is called Spatial auditory displays: Reducing cognitive load and improving driver reaction times, and it has just been published by Embedded Computing Design magazine.
 

WIRED Autopia slips into driver's seat of QNX reference vehicle

Thursday, June 14, 2012

Chances are, you've seen pictures of the new QNX reference vehicle. You may have even seen the "making of" video that QNX released a few days ago. But have you seen any video of the vehicle in action?

If not, check out this vid by Doug Newcomb of WIRED Autopia. Last week, at Telematics Detroit, Doug met up with Andrew Poliak of QNX for a tour of the vehicle and its various features, including a re-skinnable UI and voice-controlled Facebook integration. The camera was rolling, and here's what it caught:


Find me a Starbucks! QNX concept car showcases power of WATSON speech engine

Thursday, April 19, 2012

Yes, you can talk to the QNX concept car and tell it what to do.

Recently, our friends at AT&T invited us to bring the concept car to their "Living the Networked Life" event in New York. We said yes, of course! After all, what could be cooler than riding the streets of Gotham City in a digitally pimped-out Porsche 911?

Kidding aside, the event provided an excellent opportunity to demonstrate how the car takes advantage of WATSON, AT&T's natural-language speech engine. To get an idea of what WATSON can do, check out this video from Terrence O'Brien of Engadget:



For the full Engadget article, click here. And stay tuned for more updates from the Living the Networked Life event.
 

Making of the QNX concept car... honest

Tuesday, April 17, 2012

We created this video as a backdrop for CES 2012 where we unveiled the latest QNX concept car (a Porsche 911). Good thing, too, as people clearly stated that they would not have believed we did this cool retrofit ourselves without proof.


 

Video: The secret to making hands-free noise-free

Tuesday, December 6, 2011

 
Explaining a highly technical product to a broad audience is tough. To succeed, you must reach out to people on their own terms, without being condescending. Most people love a good explanation, but everyone hates being talked down to.

Case in point: The QNX Acoustic Processing Suite. This software runs in millions of cars and offers a benefit that everyone can relate to: clear, rich, easy-to-understand hands-free calls. But once you start explaining how the suite does this, it's easy to get mired in technical jargon and to forget the bigger picture — something that even a technical audience wants to see.

So we dropped the jargon and opted for a creative approach. It involves a marching band, a rock guitarist, and, for good measure, an electric fan with a really long extension cord. Seriously.

Intrigued yet? Well, then, grab some popcorn and dim the lights:




Interested in learning more about this technology? Check out the acoustic processing page on the QNX website.

BTW, companies that use the QNX Acoustic Processing Suite in their products include OnStar, whose FMV aftermarket mirror recently won a CES Innovations Design and Engineering Award.

Posted by Paul Leroux
 

New release of QNX acoustic processing suite means less noise, less tuning for hands-free systems

Tuesday, October 18, 2011

Paul Leroux
This just in: QNX has released version 2.0 of its acoustic processing suite, a modular software library designed to maximize the quality and clarity of automotive hands-free systems.

The suite, used by 18 automakers on over 100 vehicle platforms, provides modules for both the receive side and the send side of hands-free calls. The modules include acoustic echo cancellation, noise reduction, wind blocking, dynamic parametric equalization, bandwidth extension, high frequency encoding, and many others. Together, they enable high-quality voice communication, even in a noisy automotive interior.

Highlights of version 2.0 include:

Enhanced noise reduction — Minimizes audio distortions and significantly improves call clarity. Can also reconstruct speech masked by low-frequency road and engine noise.

Automatic delay calculation and compensation — Eliminates almost all product tuning, enabling automakers to save significant deployment time and expense.

Off-Axis noise rejection — Rejects sound not directly in front of a microphone or speaker, allowing dual-microphone solutions to hone in on the person speaking for greater intelligibility.

To read the press release, click here. To learn more about the acoustic processing suite, visit the QNX website.


The QNX Aviage Acoustic Processing Suite can run on the general purpose processor,
saving the cost of a DSP.


 

Wanted: Haunted Vehicles

Sunday, October 16, 2011

Halloween is just around the corner, and that reminds me of the haunted room at Lucent Bell Labs. Mind you, it wasn’t really haunted. But for a moment, I was convinced.

Let me explain. As I entered the room, I could hear two of my colleagues talking to each other, and by the sound of their voices, they were both sitting right in front of me. But when I looked, I could see only one person. Creepy, to say the least.

It took a few seconds, but I finally realized what was happening: The other colleague was in a different room, talking over a perfectly tuned prototype of a conference phone. The sense of presence was so real that I couldn’t help but feel we were all in the same room — even after I became aware of the “trick” being played!

It was then that I realized it: We don’t know what we’re missing until we experience it.

Making it real
Current telephone calls don’t sound like face-to-face conversations because the telephone network and terminals band-limit speech from about 50-10000 Hz down to 300-3400 Hz. To make matters worse, the phone’s single channel of audio eliminates spatial information about the sound source. As a result, we perceive most sounds as coming from the same point in space.

But here's the thing: The historical reasons for transmitting these single-channel narrowband speech signals no longer apply. Current technologies — such as wideband speech coders, spatial audio, and VoIP — are enabling speech communications with wider bandwidth speech and greater spatial information.

Many in the industry refer to these next-generation telecommunications systems as telepresence systems. “Telepresence” refers to the degree of realism created by a telecommunications system. Traditional systems have low telepresence while newer systems that use wider bandwidth speech and spatial audio have high telepresence.

Some people believe that a visual display is a must-have for a telepresence system. In reality, a display can decrease telepresence if its quality is poor. Experience shows that an audio-only system can have such high telepresence that people can't distinguish it from face-to-face communications — witness my haunting experience at Lucent Bell Labs.

Until recently, widespread deployment of telepresence systems has hit a roadblock: lack of standardization. Fortunately, the IETF CLUE Working Group and ITU-T Study Groups 16 and 12 are actively developing standards to remedy this situation.

Pimp my ride with telepresence
Telepresence systems have a lot to offer in an automotive environment. For instance, they could:

  • reduce driver distraction
  • make it easier to understand speech in the presence of vehicle noise
  • reduce the fatigue that comes from trying to understand a degraded voice signal

Moreover, a telepresence system makes the talker on the far end of the phone connection sound more like they are in the vehicle; it also makes the talker easier to identify.

Successful deployment of telepresence in an automotive environment depends on several factors:

  • attention to the design of vehicle platforms
  • use of high-performance acoustic processing algorithms (AEC, NR, etc.), such as those provided by the QNX acoustic processing suite
  • the ability to transport telepresence signals between telephony terminals — this is being enabled by increased VoIP availability (via LTE, for instance)

I don't know about you, but I'm looking forward to the day when my vehicle is haunted like that lab in New Jersey!

For additional reading on this topic, download the whitepaper, "Wideband Speech Communications for Automotive: the Good, the Bad, and the Ugly".

 

Total Pageviews