Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #228455 > unrolled thread
| Started by | Richard Owlett <rowlett@cloud85.net> |
|---|---|
| First post | 2020-11-04 16:00 +0100 |
| Last post | 2020-11-06 17:40 +0100 |
| Articles | 13 — 7 participants |
Back to article view | Back to linux.debian.user
Feasibility of speech recognition for note taking on dedicated laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-04 16:00 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? rhkramer@gmail.com - 2020-11-04 18:10 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? Curt <curty@free.fr> - 2020-11-04 19:00 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? rhkramer@gmail.com - 2020-11-05 00:00 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? rhkramer@gmail.com - 2020-11-05 00:10 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? <tomas@tuxteam.de> - 2020-11-05 09:10 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-05 15:30 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-05 16:30 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? deloptes <deloptes@gmail.com> - 2020-11-06 08:30 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? <tomas@tuxteam.de> - 2020-11-06 09:30 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? mick crane <mick.crane@gmail.com> - 2020-11-06 11:30 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-06 13:50 +0100
Re: Feasibility of speech recognition for note taking on dedicated laptop? Dan Hitt <dan.hitt@gmail.com> - 2020-11-06 17:40 +0100
| From | Richard Owlett <rowlett@cloud85.net> |
|---|---|
| Date | 2020-11-04 16:00 +0100 |
| Subject | Feasibility of speech recognition for note taking on dedicated laptop? |
| Message-ID | <B7sqt-98-5@gated-at.bofh.it> |
I'm a lousy typist. Trying to make notes on a laptop does not work well because typing interrupts my train of thought. Many years ago when I was a Windows user and Dragon Naturally Speaking was in its initial release I followed speech recognition casually - but not recently. Are there now end-user, Debian compatible, dictation applications that do NOT require proprietary software nor internet connectivity? My internet searching turned up primarily old material or tool-set packages packages aimed at programmers creating their own packages. Comments or suggestions? TIA
[toc] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2020-11-04 18:10 +0100 |
| Subject | Re: Feasibility of speech recognition for note taking on dedicated laptop? |
| Message-ID | <B7ush-1Di-1@gated-at.bofh.it> |
| In reply to | #228455 |
On Wednesday, November 04, 2020 09:39:44 AM Richard Owlett wrote: > I'm a lousy typist. Trying to make notes on a laptop does not work well > because typing interrupts my train of thought. > > Many years ago when I was a Windows user and Dragon Naturally Speaking > was in its initial release I followed speech recognition casually - but > not recently. > > Are there now end-user, Debian compatible, dictation applications that > do NOT require proprietary software nor internet connectivity? My > internet searching turned up primarily old material or tool-set packages > packages aimed at programmers creating their own packages. Did your Internet search turn up a package "produced" by Carnegie-Mellon? I don't remember the name of it, I'm not sure it is free, and I don't know if it has continued to be developed. I never used it, but looked into it a little back in the days when I used Dragon (a little).
[toc] | [prev] | [next] | [standalone]
| From | Curt <curty@free.fr> |
|---|---|
| Date | 2020-11-04 19:00 +0100 |
| Message-ID | <B7veG-1U6-5@gated-at.bofh.it> |
| In reply to | #228459 |
On 2020-11-04, rhkramer@gmail.com <rhkramer@gmail.com> wrote: > On Wednesday, November 04, 2020 09:39:44 AM Richard Owlett wrote: >> I'm a lousy typist. Trying to make notes on a laptop does not work well >> because typing interrupts my train of thought. >> >> Many years ago when I was a Windows user and Dragon Naturally Speaking >> was in its initial release I followed speech recognition casually - but >> not recently. >> >> Are there now end-user, Debian compatible, dictation applications that >> do NOT require proprietary software nor internet connectivity? My >> internet searching turned up primarily old material or tool-set packages >> packages aimed at programmers creating their own packages. > > Did your Internet search turn up a package "produced" by Carnegie-Mellon? > > I don't remember the name of it, I'm not sure it is free, and I don't know if > it has continued to be developed. I never used it, but looked into it a > little back in the days when I used Dragon (a little). Maybe this open source, Java (is that still a thing?) app that runs on Linux: http://www.speech.cs.cmu.edu/sphinx/dictator/
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2020-11-05 00:00 +0100 |
| Subject | Re: Feasibility of speech recognition for note taking on dedicated laptop? |
| Message-ID | <B7zUZ-4Mk-1@gated-at.bofh.it> |
| In reply to | #228461 |
On Wednesday, November 04, 2020 12:36:51 PM Curt wrote: > Maybe this open source, Java (is that still a thing?) app that runs > on Linux: > > http://www.speech.cs.cmu.edu/sphinx/dictator/ Yes, I believe that it is it, but maybe I saw an earlier version (although the web page listed above is copyrighted something to 2006). One of the things that made me uncomfortable was that it was written in Java, and I was concerned about the performance. But, I never did try it. Looks like it is pretty much dead -- I tried to access the TWiki but was denied access.
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2020-11-05 00:10 +0100 |
| Subject | Re: Feasibility of speech recognition for note taking on dedicated laptop? |
| Message-ID | <B7A4F-556-3@gated-at.bofh.it> |
| In reply to | #228469 |
On Wednesday, November 04, 2020 05:58:25 PM rhkramer@gmail.com wrote: > Looks like it is pretty much dead -- I tried to access the TWiki but was > denied access. Oh, but maybe it is more alive than I thought -- quoting from https://cmusphinx.github.io/ <quote> Oct 23, 2019 Update on CMUSphinx Project Dear users, you've might been asking yourself why there were not so many updates on CMUSphinx recently. Time goes really fast and many things change in ASR. Deep learning, huge NLP models like BERT, Tacotron and Wavenet/Waveglow/WaveRNN, Pytorch vs Tensorflow, huge datsets, chatbots and so on and so forth. Many new toolkits appear and some disappear - Eesen, Espresso, Kaldi, Wav2letter, NeMo. The whole area is thriving. CMUSphinx team has been actively participating in all those activities, creating new models, applications, helping newcomers and showing the best way to implement speech recognition system. We are here to suggest you the easiest way to start such an exciting world of speech recognition. Lately we implemented a Kaldi on Android, providing much better accuracy for large vocabulary decoding, which was hard to imagine before. If you are interested in learning more, check Alpha Cephei website, our Github and join us on Telegram and Reddit. Stay tuned! </quote>
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2020-11-05 09:10 +0100 |
| Message-ID | <B7Ivg-1XG-7@gated-at.bofh.it> |
| In reply to | #228469 |
[Multipart message — attachments visible in raw view] — view raw
On Wed, Nov 04, 2020 at 05:58:25PM -0500, rhkramer@gmail.com wrote:
> On Wednesday, November 04, 2020 12:36:51 PM Curt wrote:
> > Maybe this open source, Java (is that still a thing?) app that runs
> > on Linux:
> >
> > http://www.speech.cs.cmu.edu/sphinx/dictator/
>
> Yes, I believe that it is it, but maybe I saw an earlier version (although the
> web page listed above is copyrighted something to 2006).
>
> One of the things that made me uncomfortable was that it was written in Java,
> and I was concerned about the performance. But, I never did try it.
>
> Looks like it is pretty much dead -- I tried to access the TWiki but was
> denied access.
A more recent project seems to be Mozilla Foundation's DeepSpeech [1]
"DeepSpeech is an open source embedded (offline, on-device)
speech-to-text engine which can run in real time on devices
ranging from a Raspberry Pi 4 to high power GPU servers."
(Sorry for linking to Github. OTOH, they seem to have some page for
this, but it's a Javascript-only white hole [2], so I don't know
what's in there)
[1] https://github.com/mozilla/DeepSpeech
[2] https://commonvoice.mozilla.org/
- t
[toc] | [prev] | [next] | [standalone]
| From | Richard Owlett <rowlett@cloud85.net> |
|---|---|
| Date | 2020-11-05 15:30 +0100 |
| Message-ID | <B7OqZ-5OF-3@gated-at.bofh.it> |
| In reply to | #228477 |
On 11/05/2020 02:08 AM, tomas@tuxteam.de wrote: > On Wed, Nov 04, 2020 at 05:58:25PM -0500, rhkramer@gmail.com wrote: >> On Wednesday, November 04, 2020 12:36:51 PM Curt wrote: >>> Maybe this open source, Java (is that still a thing?) app that runs >>> on Linux: >>> >>> http://www.speech.cs.cmu.edu/sphinx/dictator/ >> >> Yes, I believe that it is it, but maybe I saw an earlier version (although the >> web page listed above is copyrighted something to 2006). >> >> One of the things that made me uncomfortable was that it was written in Java, >> and I was concerned about the performance. But, I never did try it. >> >> Looks like it is pretty much dead -- I tried to access the TWiki but was >> denied access. > > A more recent project seems to be Mozilla Foundation's DeepSpeech [1] > > "DeepSpeech is an open source embedded (offline, on-device) > speech-to-text engine which can run in real time on devices > ranging from a Raspberry Pi 4 to high power GPU servers." > > (Sorry for linking to Github. OTOH, they seem to have some page for > this, but it's a Javascript-only white hole [2], so I don't know > what's in there) > > [1] https://github.com/mozilla/DeepSpeech > [2] https://commonvoice.mozilla.org/ > > - t > My impression of the CMU project(s) is that the focus is more on developers of speech enabled software than end-users of the application. Initial browsing indicates DeepSpeech will likely be more appropriate. Some links from [1] state a requirement for JavaScript but display a blank screen even when JavaScript has been enabled. The same is true of [2] itself. I suspect "browser sniffing" as [https://chat.mozilla.org/#/room/#machinelearning:mozilla.org], pointed to by [https://github.com/mozilla/DeepSpeech/blob/master/SUPPORT.rst], explicitly states: > > Your browser can't run Element > > Element uses many advanced browser features, some of which are not available or experimental in your current browser. > > Please install Chrome, Firefox, or Safari for the best experience. > Use Element on mobile My current browser is SeaMonkey 2.49.4 [cookies disabled] running on Debian 9. I intend to do an install of Debian 10 on another laptop in the next week an will install current Firefox to see if that is the only problem. I will also try the public machines at the local library as a double-check. It may be out of date information, but [https://github.com/mozilla/DeepSpeech/wiki#why-cant-i-speak-directly-to-deepspeech-instead-of-first-making-an-audio-recording] says: > Why can't I speak directly to DeepSpeech instead of first making an audio recording? > > We are providing inference tools as a way to easily test the system, but building > upon that is open to anyone. Having to deal with more interactive UX is out of the > scope of the current target of those tools. One of the links I went to explicitly referred to doing "real time" speech recognition [it may have been written by an application developer]. More later. Thank you.
[toc] | [prev] | [next] | [standalone]
| From | Richard Owlett <rowlett@cloud85.net> |
|---|---|
| Date | 2020-11-05 16:30 +0100 |
| Message-ID | <B7Pn4-6qx-3@gated-at.bofh.it> |
| In reply to | #228486 |
On 11/05/2020 08:20 AM, Richard Owlett wrote: > On 11/05/2020 02:08 AM, tomas@tuxteam.de wrote: >> On Wed, Nov 04, 2020 at 05:58:25PM -0500, rhkramer@gmail.com wrote: >>> On Wednesday, November 04, 2020 12:36:51 PM Curt wrote: >>>> Maybe this open source, Java (is that still a thing?) app that runs >>>> on Linux: >>>> >>>> http://www.speech.cs.cmu.edu/sphinx/dictator/ >>> >>> Yes, I believe that it is it, but maybe I saw an earlier version >>> (although the >>> web page listed above is copyrighted something to 2006). >>> >>> One of the things that made me uncomfortable was that it was written >>> in Java, >>> and I was concerned about the performance. But, I never did try it. >>> >>> Looks like it is pretty much dead -- I tried to access the TWiki but was >>> denied access. >> >> A more recent project seems to be Mozilla Foundation's DeepSpeech [1] >> >> "DeepSpeech is an open source embedded (offline, on-device) >> speech-to-text engine which can run in real time on devices >> ranging from a Raspberry Pi 4 to high power GPU servers." >> >> (Sorry for linking to Github. OTOH, they seem to have some page for >> this, but it's a Javascript-only white hole [2], so I don't know >> what's in there) >> >> [1] https://github.com/mozilla/DeepSpeech >> [2] https://commonvoice.mozilla.org/ >> >> - t >> > > My impression of the CMU project(s) is that the focus is more on > developers of speech enabled software than end-users of the application. > > Initial browsing indicates DeepSpeech will likely be more appropriate. > > Some links from [1] state a requirement for JavaScript but display a > blank screen even when JavaScript has been enabled. The same is true of > [2] itself. > > I suspect "browser sniffing" as > [https://chat.mozilla.org/#/room/#machinelearning:mozilla.org], pointed > to by [https://github.com/mozilla/DeepSpeech/blob/master/SUPPORT.rst], > explicitly states: >> >> Your browser can't run Element >> >> Element uses many advanced browser features, some of which are not >> available or experimental in your current browser. >> >> Please install Chrome, Firefox, or Safari for the best experience. >> Use Element on mobile [ADDENDUM] One of my laptops has Firefox 60 on Debian 10 with JavaScript and cookies enabled. [2] displays properly. [https://chat.mozilla.org/#/room/#machinelearning:mozilla.org] will apparently display properly after complaining that the browser was not current version. > > My current browser is SeaMonkey 2.49.4 [cookies disabled] running on > Debian 9. I intend to do an install of Debian 10 on another laptop in > the next week an will install current Firefox to see if that is the only > problem. I will also try the public machines at the local library as a > double-check. > > It may be out of date information, but > [https://github.com/mozilla/DeepSpeech/wiki#why-cant-i-speak-directly-to-deepspeech-instead-of-first-making-an-audio-recording] > says: >> Why can't I speak directly to DeepSpeech instead of first making an >> audio recording? >> >> We are providing inference tools as a way to easily test the system, >> but building >> upon that is open to anyone. Having to deal with more interactive UX >> is out of the >> scope of the current target of those tools. > One of the links I went to explicitly referred to doing "real time" > speech recognition [it may have been written by an application developer]. > > More later. > Thank you. > > > > > > >
[toc] | [prev] | [next] | [standalone]
| From | deloptes <deloptes@gmail.com> |
|---|---|
| Date | 2020-11-06 08:30 +0100 |
| Subject | Re: Feasibility of speech recognition for note taking on dedicated laptop? |
| Message-ID | <B84m5-7sf-3@gated-at.bofh.it> |
| In reply to | #228455 |
Richard Owlett wrote: > I'm a lousy typist. Trying to make notes on a laptop does not work well > because typing interrupts my train of thought. > > Many years ago when I was a Windows user and Dragon Naturally Speaking > was in its initial release I followed speech recognition casually - but > not recently. > > Are there now end-user, Debian compatible, dictation applications that > do NOT require proprietary software nor internet connectivity? My > internet searching turned up primarily old material or tool-set packages > packages aimed at programmers creating their own packages. > > Comments or suggestions? Again one of these topics, where people post about software they do not actually use. Let me comment here my impressions. I studied speech processing and wrote my thesis on dialog systems in 2007. Until about 2005 there were still some open source tools like ViaVoice by IBM. Basically all of this was dropped by 2010 - no idea why - might be something related to Google/Amazon, costs, patents or whatever else. I have not heard or seen any useful Linux tool - I mean not something like Alexa that would send the audio recorded on the device to NSA for processing. One of the problems is surely the complexity and hardware requirements to run such tool (I mean a dialog engine like Alexa). However I do not understand why the linux community does not have a usable STT application as even 10+y ago there were usable tools for windows. And here we come to the original question. I have seen and tested some of the windows tools 10+y ago. They implement some type of learning, adapting themselves to the way you speak etc. In 2-3 months period the success rate goes to 95-98%. Price back then was affordable - AFAIR 200-400 US My advise, if you need such tool - look for commercial applications and forget the linux crap. I am willing to convince myself in the opposite, but looking what happens on the linux scene ... I've lost faith
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2020-11-06 09:30 +0100 |
| Message-ID | <B85ia-82j-5@gated-at.bofh.it> |
| In reply to | #228496 |
[Multipart message — attachments visible in raw view] — view raw
On Fri, Nov 06, 2020 at 08:25:36AM +0100, deloptes wrote: [...] > Again one of these topics, where people post about software they do not > actually use. In my case, that's true. I do follow the topic, but from some safe distance. > Let me comment here my impressions. I studied speech processing and wrote my > thesis on dialog systems in 2007. Until about 2005 there were still some > open source tools like ViaVoice by IBM. Basically all of this was dropped > by 2010 - no idea why - might be something related to Google/Amazon, costs, > patents or whatever else. AFAIU, things have changed since them. Much more machine learing is involved these days, due to the availability of huge parallel computing power even on mobile devices (GPU, dedicated hardware -- the marketing buzzword is "AI edge"). The challenge these days is in collecting enough and diverse sample data to develop (aka "deep learn") a suitable "model". And that's where Mozilla's once-deep pockets tried to help. But, as I said, I'm not an expert, but I play one on TV. Cheers - t
[toc] | [prev] | [next] | [standalone]
| From | mick crane <mick.crane@gmail.com> |
|---|---|
| Date | 2020-11-06 11:30 +0100 |
| Message-ID | <B87ai-NP-5@gated-at.bofh.it> |
| In reply to | #228496 |
On 2020-11-06 07:25, deloptes wrote: > Let me comment here my impressions. I studied speech processing and > wrote my > thesis on dialog systems in 2007. Until about 2005 there were still > some > open source tools like ViaVoice by IBM. Basically all of this was > dropped > by 2010 - no idea why - might be something related to Google/Amazon, > costs, > patents or whatever else. > In like 1980 or something noticed voice recognition chip in radio shack for not very much money. Although curious at that time was, and still is, way, way beyond my understanding of what you might do with it. mick -- Key ID 4BFEBB31
[toc] | [prev] | [next] | [standalone]
| From | Richard Owlett <rowlett@cloud85.net> |
|---|---|
| Date | 2020-11-06 13:50 +0100 |
| Message-ID | <B89lL-237-5@gated-at.bofh.it> |
| In reply to | #228496 |
On 11/06/2020 01:25 AM, deloptes wrote: > Richard Owlett wrote: > >> I'm a lousy typist. Trying to make notes on a laptop does not work well >> because typing interrupts my train of thought. >> >> Many years ago when I was a Windows user and Dragon Naturally Speaking >> was in its initial release I followed speech recognition casually - but >> not recently. >> >> Are there now end-user, Debian compatible, dictation applications that >> do NOT require proprietary software nor internet connectivity? My >> internet searching turned up primarily old material or tool-set packages >> packages aimed at programmers creating their own packages. >> >> Comments or suggestions? > > Again one of these topics, where people post about software they do not > actually use. > > Let me comment here my impressions. I studied speech processing and wrote my > thesis on dialog systems in 2007. Until about 2005 there were still some > open source tools like ViaVoice by IBM. Basically all of this was dropped > by 2010 - no idea why - might be something related to Google/Amazon, costs, > patents or whatever else. > > I have not heard or seen any useful Linux tool - I mean not something like > Alexa that would send the audio recorded on the device to NSA for > processing. > One of the problems is surely the complexity and hardware requirements to > run such tool (I mean a dialog engine like Alexa). > > However I do not understand why the linux community does not have a usable > STT application as even 10+y ago there were usable tools for windows. > And here we come to the original question. I have seen and tested some of > the windows tools 10+y ago. They implement some type of learning, adapting > themselves to the way you speak etc. In 2-3 months period the success rate > goes to 95-98%. Price back then was affordable - AFAIR 200-400 US > My advise, if you need such tool - look for commercial applications and > forget the linux crap. I am willing to convince myself in the opposite, but > looking what happens on the linux scene ... I've lost faith > I abandoned proprietary software when Windows XP Pro SP3 was current. Vendors think the user should only be interested in problems for which they have a $olution. I have an absolute requirement that any software I acquire run under Linux - preferably Debian. One focus of DeepSpeech seems to be aimed in my direction. It may not be as much of an "out of the box" solution as I would prefer. But it seems close enough to merit more investigation.
[toc] | [prev] | [next] | [standalone]
| From | Dan Hitt <dan.hitt@gmail.com> |
|---|---|
| Date | 2020-11-06 17:40 +0100 |
| Subject | Re: Feasibility of speech recognition for note taking on dedicated laptop? |
| Message-ID | <B8cWl-4e8-1@gated-at.bofh.it> |
| In reply to | #228496 |
[Multipart message — attachments visible in raw view] — view raw
On Thu, Nov 5, 2020 at 11:26 PM deloptes <deloptes@gmail.com> wrote: > Richard Owlett wrote: > > ..... > > Are there now end-user, Debian compatible, dictation applications that > > do NOT require proprietary software nor internet connectivity? My > > internet searching turned up primarily old material or tool-set packages > > packages aimed at programmers creating their own packages. > > > > Comments or suggestions? > > Again one of these topics, where people post about software they do not > actually use. > > Let me comment here my impressions. I studied speech processing and wrote > my > thesis on dialog systems in 2007. Until about 2005 there were still some > open source tools like ViaVoice by IBM. Basically all of this was dropped > by 2010 - no idea why - might be something related to Google/Amazon, costs, > patents or whatever else. > > > > > Hi Deloptes, If your thesis is available online, could you please post a link to it? TIA! dan
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.user
csiph-web