Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #228455 > unrolled thread

Feasibility of speech recognition for note taking on dedicated laptop?

Started byRichard Owlett <rowlett@cloud85.net>
First post2020-11-04 16:00 +0100
Last post2020-11-06 17:40 +0100
Articles 13 — 7 participants

Back to article view | Back to linux.debian.user


Contents

  Feasibility of speech recognition for note taking on dedicated  laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-04 16:00 +0100
    Re: Feasibility of speech recognition for note taking on dedicated laptop? rhkramer@gmail.com - 2020-11-04 18:10 +0100
      Re: Feasibility of speech recognition for note taking on dedicated  laptop? Curt <curty@free.fr> - 2020-11-04 19:00 +0100
        Re: Feasibility of speech recognition for note taking on dedicated laptop? rhkramer@gmail.com - 2020-11-05 00:00 +0100
          Re: Feasibility of speech recognition for note taking on dedicated laptop? rhkramer@gmail.com - 2020-11-05 00:10 +0100
          Re: Feasibility of speech recognition for note taking on dedicated  laptop? <tomas@tuxteam.de> - 2020-11-05 09:10 +0100
            Re: Feasibility of speech recognition for note taking on dedicated  laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-05 15:30 +0100
              Re: Feasibility of speech recognition for note taking on dedicated  laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-05 16:30 +0100
    Re: Feasibility of speech recognition for note taking on dedicated laptop? deloptes <deloptes@gmail.com> - 2020-11-06 08:30 +0100
      Re: Feasibility of speech recognition for note taking on dedicated  laptop? <tomas@tuxteam.de> - 2020-11-06 09:30 +0100
      Re: Feasibility of speech recognition for note taking on dedicated  laptop? mick crane <mick.crane@gmail.com> - 2020-11-06 11:30 +0100
      Re: Feasibility of speech recognition for note taking on dedicated  laptop? Richard Owlett <rowlett@cloud85.net> - 2020-11-06 13:50 +0100
      Re: Feasibility of speech recognition for note taking on dedicated laptop? Dan Hitt <dan.hitt@gmail.com> - 2020-11-06 17:40 +0100

#228455 — Feasibility of speech recognition for note taking on dedicated laptop?

FromRichard Owlett <rowlett@cloud85.net>
Date2020-11-04 16:00 +0100
SubjectFeasibility of speech recognition for note taking on dedicated laptop?
Message-ID<B7sqt-98-5@gated-at.bofh.it>
I'm a lousy typist. Trying to make notes on a laptop does not work well 
because typing interrupts my train of thought.

Many years ago when I was a Windows user and Dragon Naturally Speaking 
was in its initial release I followed speech recognition casually - but 
not recently.

Are there now end-user, Debian compatible, dictation applications that 
do NOT require proprietary software nor internet connectivity? My 
internet searching turned up primarily old material or tool-set packages 
packages aimed at programmers creating their own packages.

Comments or suggestions?
TIA

[toc] | [next] | [standalone]


#228459 — Re: Feasibility of speech recognition for note taking on dedicated laptop?

Fromrhkramer@gmail.com
Date2020-11-04 18:10 +0100
SubjectRe: Feasibility of speech recognition for note taking on dedicated laptop?
Message-ID<B7ush-1Di-1@gated-at.bofh.it>
In reply to#228455
On Wednesday, November 04, 2020 09:39:44 AM Richard Owlett wrote:
> I'm a lousy typist. Trying to make notes on a laptop does not work well
> because typing interrupts my train of thought.
> 
> Many years ago when I was a Windows user and Dragon Naturally Speaking
> was in its initial release I followed speech recognition casually - but
> not recently.
> 
> Are there now end-user, Debian compatible, dictation applications that
> do NOT require proprietary software nor internet connectivity? My
> internet searching turned up primarily old material or tool-set packages
> packages aimed at programmers creating their own packages.

Did your Internet search turn up a package "produced" by Carnegie-Mellon?

I don't remember the name of it, I'm not sure it is free, and I don't know if 
it has continued to be developed.  I never used it, but looked into it a 
little back in the days when I used Dragon (a little).

[toc] | [prev] | [next] | [standalone]


#228461

FromCurt <curty@free.fr>
Date2020-11-04 19:00 +0100
Message-ID<B7veG-1U6-5@gated-at.bofh.it>
In reply to#228459
On 2020-11-04, rhkramer@gmail.com <rhkramer@gmail.com> wrote:
> On Wednesday, November 04, 2020 09:39:44 AM Richard Owlett wrote:
>> I'm a lousy typist. Trying to make notes on a laptop does not work well
>> because typing interrupts my train of thought.
>> 
>> Many years ago when I was a Windows user and Dragon Naturally Speaking
>> was in its initial release I followed speech recognition casually - but
>> not recently.
>> 
>> Are there now end-user, Debian compatible, dictation applications that
>> do NOT require proprietary software nor internet connectivity? My
>> internet searching turned up primarily old material or tool-set packages
>> packages aimed at programmers creating their own packages.
>
> Did your Internet search turn up a package "produced" by Carnegie-Mellon?
>
> I don't remember the name of it, I'm not sure it is free, and I don't know if 
> it has continued to be developed.  I never used it, but looked into it a 
> little back in the days when I used Dragon (a little).

Maybe this open source, Java (is that still a thing?) app that runs
on Linux:

http://www.speech.cs.cmu.edu/sphinx/dictator/

[toc] | [prev] | [next] | [standalone]


#228469 — Re: Feasibility of speech recognition for note taking on dedicated laptop?

Fromrhkramer@gmail.com
Date2020-11-05 00:00 +0100
SubjectRe: Feasibility of speech recognition for note taking on dedicated laptop?
Message-ID<B7zUZ-4Mk-1@gated-at.bofh.it>
In reply to#228461
On Wednesday, November 04, 2020 12:36:51 PM Curt wrote:
> Maybe this open source, Java (is that still a thing?) app that runs
> on Linux:
> 
> http://www.speech.cs.cmu.edu/sphinx/dictator/

Yes, I believe that it is it, but maybe I saw an earlier version (although the 
web page listed above is copyrighted something to 2006).

One of the things that made me uncomfortable was that it was written in Java, 
and I was concerned about the performance.  But, I never did try it.

Looks like it is pretty much dead -- I tried to access the TWiki but was 
denied access.

[toc] | [prev] | [next] | [standalone]


#228470 — Re: Feasibility of speech recognition for note taking on dedicated laptop?

Fromrhkramer@gmail.com
Date2020-11-05 00:10 +0100
SubjectRe: Feasibility of speech recognition for note taking on dedicated laptop?
Message-ID<B7A4F-556-3@gated-at.bofh.it>
In reply to#228469
On Wednesday, November 04, 2020 05:58:25 PM rhkramer@gmail.com wrote:
> Looks like it is pretty much dead -- I tried to access the TWiki but was
> denied access.

Oh, but maybe it is more alive than I thought -- quoting from 
https://cmusphinx.github.io/

<quote>
Oct 23, 2019
Update on CMUSphinx Project

Dear users, you've might been asking yourself why there were not so many 
updates on CMUSphinx recently. Time goes really fast and many things change in 
ASR. Deep learning, huge NLP models like BERT, Tacotron and 
Wavenet/Waveglow/WaveRNN, Pytorch vs Tensorflow, huge datsets, chatbots and so 
on and so forth. Many new toolkits appear and some disappear - Eesen, 
Espresso, Kaldi, Wav2letter, NeMo. The whole area is thriving.

CMUSphinx team has been actively participating in all those activities, 
creating new models, applications, helping newcomers and showing the best way 
to implement speech recognition system. We are here to suggest you the easiest 
way to start such an exciting world of speech recognition. Lately we 
implemented a Kaldi on Android, providing much better accuracy for large 
vocabulary decoding, which was hard to imagine before.

If you are interested in learning more, check Alpha Cephei website, our Github 
and join us on Telegram and Reddit.

Stay tuned!
</quote>

[toc] | [prev] | [next] | [standalone]


#228477

From<tomas@tuxteam.de>
Date2020-11-05 09:10 +0100
Message-ID<B7Ivg-1XG-7@gated-at.bofh.it>
In reply to#228469

[Multipart message — attachments visible in raw view] — view raw

On Wed, Nov 04, 2020 at 05:58:25PM -0500, rhkramer@gmail.com wrote:
> On Wednesday, November 04, 2020 12:36:51 PM Curt wrote:
> > Maybe this open source, Java (is that still a thing?) app that runs
> > on Linux:
> > 
> > http://www.speech.cs.cmu.edu/sphinx/dictator/
> 
> Yes, I believe that it is it, but maybe I saw an earlier version (although the 
> web page listed above is copyrighted something to 2006).
> 
> One of the things that made me uncomfortable was that it was written in Java, 
> and I was concerned about the performance.  But, I never did try it.
> 
> Looks like it is pretty much dead -- I tried to access the TWiki but was 
> denied access.

A more recent project seems to be Mozilla Foundation's DeepSpeech [1]

    "DeepSpeech is an open source embedded (offline, on-device)
     speech-to-text engine which can run in real time on devices
     ranging from a Raspberry Pi 4 to high power GPU servers."

(Sorry for linking to Github. OTOH, they seem to have some page for
this, but it's a Javascript-only white hole [2], so I don't know
what's in there)

[1] https://github.com/mozilla/DeepSpeech
[2] https://commonvoice.mozilla.org/

 - t

[toc] | [prev] | [next] | [standalone]


#228486

FromRichard Owlett <rowlett@cloud85.net>
Date2020-11-05 15:30 +0100
Message-ID<B7OqZ-5OF-3@gated-at.bofh.it>
In reply to#228477
On 11/05/2020 02:08 AM, tomas@tuxteam.de wrote:
> On Wed, Nov 04, 2020 at 05:58:25PM -0500, rhkramer@gmail.com wrote:
>> On Wednesday, November 04, 2020 12:36:51 PM Curt wrote:
>>> Maybe this open source, Java (is that still a thing?) app that runs
>>> on Linux:
>>>
>>> http://www.speech.cs.cmu.edu/sphinx/dictator/
>>
>> Yes, I believe that it is it, but maybe I saw an earlier version (although the
>> web page listed above is copyrighted something to 2006).
>>
>> One of the things that made me uncomfortable was that it was written in Java,
>> and I was concerned about the performance.  But, I never did try it.
>>
>> Looks like it is pretty much dead -- I tried to access the TWiki but was
>> denied access.
> 
> A more recent project seems to be Mozilla Foundation's DeepSpeech [1]
> 
>      "DeepSpeech is an open source embedded (offline, on-device)
>       speech-to-text engine which can run in real time on devices
>       ranging from a Raspberry Pi 4 to high power GPU servers."
> 
> (Sorry for linking to Github. OTOH, they seem to have some page for
> this, but it's a Javascript-only white hole [2], so I don't know
> what's in there)
> 
> [1] https://github.com/mozilla/DeepSpeech
> [2] https://commonvoice.mozilla.org/
> 
>   - t
> 

My impression of the CMU project(s) is that the focus is more on 
developers of speech enabled software than end-users of the application.

Initial browsing indicates DeepSpeech will likely be more appropriate.

Some links from [1] state a requirement for JavaScript but display a 
blank screen even when JavaScript has been enabled. The same is true of 
[2] itself.

I suspect "browser sniffing" as 
[https://chat.mozilla.org/#/room/#machinelearning:mozilla.org], pointed 
to by [https://github.com/mozilla/DeepSpeech/blob/master/SUPPORT.rst], 
explicitly states:
> 
> Your browser can't run Element
> 
> Element uses many advanced browser features, some of which are not available or experimental in your current browser.
> 
> Please install Chrome, Firefox, or Safari for the best experience.
> Use Element on mobile

My current browser is SeaMonkey 2.49.4 [cookies disabled] running on 
Debian 9. I intend to do an install of Debian 10 on another laptop in 
the next week an will install current Firefox to see if that is the only 
problem. I will also try the public machines at the local library as a 
double-check.

It may be out of date information, but 
[https://github.com/mozilla/DeepSpeech/wiki#why-cant-i-speak-directly-to-deepspeech-instead-of-first-making-an-audio-recording] 
says:
> Why can't I speak directly to DeepSpeech instead of first making an audio recording?
> 
> We are providing inference tools as a way to easily test the system, but building
> upon that is open to anyone. Having to deal with more interactive UX is out of the
> scope of the current target of those tools.
One of the links I went to explicitly referred to doing "real time" 
speech recognition [it may have been written by an application developer].

More later.
Thank you.

[toc] | [prev] | [next] | [standalone]


#228487

FromRichard Owlett <rowlett@cloud85.net>
Date2020-11-05 16:30 +0100
Message-ID<B7Pn4-6qx-3@gated-at.bofh.it>
In reply to#228486
On 11/05/2020 08:20 AM, Richard Owlett wrote:
> On 11/05/2020 02:08 AM, tomas@tuxteam.de wrote:
>> On Wed, Nov 04, 2020 at 05:58:25PM -0500, rhkramer@gmail.com wrote:
>>> On Wednesday, November 04, 2020 12:36:51 PM Curt wrote:
>>>> Maybe this open source, Java (is that still a thing?) app that runs
>>>> on Linux:
>>>>
>>>> http://www.speech.cs.cmu.edu/sphinx/dictator/
>>>
>>> Yes, I believe that it is it, but maybe I saw an earlier version 
>>> (although the
>>> web page listed above is copyrighted something to 2006).
>>>
>>> One of the things that made me uncomfortable was that it was written 
>>> in Java,
>>> and I was concerned about the performance.  But, I never did try it.
>>>
>>> Looks like it is pretty much dead -- I tried to access the TWiki but was
>>> denied access.
>>
>> A more recent project seems to be Mozilla Foundation's DeepSpeech [1]
>>
>>      "DeepSpeech is an open source embedded (offline, on-device)
>>       speech-to-text engine which can run in real time on devices
>>       ranging from a Raspberry Pi 4 to high power GPU servers."
>>
>> (Sorry for linking to Github. OTOH, they seem to have some page for
>> this, but it's a Javascript-only white hole [2], so I don't know
>> what's in there)
>>
>> [1] https://github.com/mozilla/DeepSpeech
>> [2] https://commonvoice.mozilla.org/
>>
>>   - t
>>
> 
> My impression of the CMU project(s) is that the focus is more on 
> developers of speech enabled software than end-users of the application.
> 
> Initial browsing indicates DeepSpeech will likely be more appropriate.
> 
> Some links from [1] state a requirement for JavaScript but display a 
> blank screen even when JavaScript has been enabled. The same is true of 
> [2] itself.
> 
> I suspect "browser sniffing" as 
> [https://chat.mozilla.org/#/room/#machinelearning:mozilla.org], pointed 
> to by [https://github.com/mozilla/DeepSpeech/blob/master/SUPPORT.rst], 
> explicitly states:
>>
>> Your browser can't run Element
>>
>> Element uses many advanced browser features, some of which are not 
>> available or experimental in your current browser.
>>
>> Please install Chrome, Firefox, or Safari for the best experience.
>> Use Element on mobile

[ADDENDUM]
One of my laptops has Firefox 60 on Debian 10 with JavaScript and 
cookies enabled.
[2] displays properly.
[https://chat.mozilla.org/#/room/#machinelearning:mozilla.org] will 
apparently display properly after complaining that the browser was not 
current version.



> 
> My current browser is SeaMonkey 2.49.4 [cookies disabled] running on 
> Debian 9. I intend to do an install of Debian 10 on another laptop in 
> the next week an will install current Firefox to see if that is the only 
> problem. I will also try the public machines at the local library as a 
> double-check.
> 
> It may be out of date information, but 
> [https://github.com/mozilla/DeepSpeech/wiki#why-cant-i-speak-directly-to-deepspeech-instead-of-first-making-an-audio-recording] 
> says:
>> Why can't I speak directly to DeepSpeech instead of first making an 
>> audio recording?
>>
>> We are providing inference tools as a way to easily test the system, 
>> but building
>> upon that is open to anyone. Having to deal with more interactive UX 
>> is out of the
>> scope of the current target of those tools.
> One of the links I went to explicitly referred to doing "real time" 
> speech recognition [it may have been written by an application developer].
> 
> More later.
> Thank you.
> 
> 
> 
> 
> 
> 
> 

[toc] | [prev] | [next] | [standalone]


#228496 — Re: Feasibility of speech recognition for note taking on dedicated laptop?

Fromdeloptes <deloptes@gmail.com>
Date2020-11-06 08:30 +0100
SubjectRe: Feasibility of speech recognition for note taking on dedicated laptop?
Message-ID<B84m5-7sf-3@gated-at.bofh.it>
In reply to#228455
Richard Owlett wrote:

> I'm a lousy typist. Trying to make notes on a laptop does not work well
> because typing interrupts my train of thought.
> 
> Many years ago when I was a Windows user and Dragon Naturally Speaking
> was in its initial release I followed speech recognition casually - but
> not recently.
> 
> Are there now end-user, Debian compatible, dictation applications that
> do NOT require proprietary software nor internet connectivity? My
> internet searching turned up primarily old material or tool-set packages
> packages aimed at programmers creating their own packages.
> 
> Comments or suggestions?

Again one of these topics, where people post about software they do not
actually use.

Let me comment here my impressions. I studied speech processing and wrote my
thesis on dialog systems in 2007. Until about 2005 there were still some
open source tools like ViaVoice by IBM. Basically all of this was dropped
by 2010 - no idea why - might be something related to Google/Amazon, costs,
patents or whatever else.

I have not heard or seen any useful Linux tool - I mean not something like
Alexa that would send the audio recorded on the device to NSA for
processing.
One of the problems is surely the complexity and hardware requirements to
run such tool (I mean a dialog engine like Alexa).

However I do not understand why the linux community does not have a usable
STT application as even 10+y ago there were usable tools for windows.
And here we come to the original question. I have seen and tested some of
the windows tools 10+y ago. They implement some type of learning, adapting
themselves to the way you speak etc. In 2-3 months period the success rate
goes to 95-98%. Price back then was affordable - AFAIR 200-400 US
My advise, if you need such tool - look for commercial applications and
forget the linux crap. I am willing to convince myself in the opposite, but
looking what happens on the linux scene ... I've lost faith 

[toc] | [prev] | [next] | [standalone]


#228498

From<tomas@tuxteam.de>
Date2020-11-06 09:30 +0100
Message-ID<B85ia-82j-5@gated-at.bofh.it>
In reply to#228496

[Multipart message — attachments visible in raw view] — view raw

On Fri, Nov 06, 2020 at 08:25:36AM +0100, deloptes wrote:

[...]

> Again one of these topics, where people post about software they do not
> actually use.

In my case, that's true. I do follow the topic, but from some
safe distance.

> Let me comment here my impressions. I studied speech processing and wrote my
> thesis on dialog systems in 2007. Until about 2005 there were still some
> open source tools like ViaVoice by IBM. Basically all of this was dropped
> by 2010 - no idea why - might be something related to Google/Amazon, costs,
> patents or whatever else.

AFAIU, things have changed since them. Much more machine learing is
involved these days, due to the availability of huge parallel computing
power even on mobile devices (GPU, dedicated hardware -- the marketing
buzzword is "AI edge").

The challenge these days is in collecting enough and diverse sample
data to develop (aka "deep learn") a suitable "model". And that's where
Mozilla's once-deep pockets tried to help.

But, as I said, I'm not an expert, but I play one on TV.

Cheers
 - t

[toc] | [prev] | [next] | [standalone]


#228500

Frommick crane <mick.crane@gmail.com>
Date2020-11-06 11:30 +0100
Message-ID<B87ai-NP-5@gated-at.bofh.it>
In reply to#228496
On 2020-11-06 07:25, deloptes wrote:

> Let me comment here my impressions. I studied speech processing and 
> wrote my
> thesis on dialog systems in 2007. Until about 2005 there were still 
> some
> open source tools like ViaVoice by IBM. Basically all of this was 
> dropped
> by 2010 - no idea why - might be something related to Google/Amazon, 
> costs,
> patents or whatever else.
> 
In like 1980 or something noticed voice recognition chip in radio shack 
for not very much money.
Although curious at that time was, and still is, way, way beyond my 
understanding of what you might do with it.

mick
-- 
Key ID    4BFEBB31

[toc] | [prev] | [next] | [standalone]


#228504

FromRichard Owlett <rowlett@cloud85.net>
Date2020-11-06 13:50 +0100
Message-ID<B89lL-237-5@gated-at.bofh.it>
In reply to#228496
On 11/06/2020 01:25 AM, deloptes wrote:
> Richard Owlett wrote:
> 
>> I'm a lousy typist. Trying to make notes on a laptop does not work well
>> because typing interrupts my train of thought.
>>
>> Many years ago when I was a Windows user and Dragon Naturally Speaking
>> was in its initial release I followed speech recognition casually - but
>> not recently.
>>
>> Are there now end-user, Debian compatible, dictation applications that
>> do NOT require proprietary software nor internet connectivity? My
>> internet searching turned up primarily old material or tool-set packages
>> packages aimed at programmers creating their own packages.
>>
>> Comments or suggestions?
> 
> Again one of these topics, where people post about software they do not
> actually use.
> 
> Let me comment here my impressions. I studied speech processing and wrote my
> thesis on dialog systems in 2007. Until about 2005 there were still some
> open source tools like ViaVoice by IBM. Basically all of this was dropped
> by 2010 - no idea why - might be something related to Google/Amazon, costs,
> patents or whatever else.
> 
> I have not heard or seen any useful Linux tool - I mean not something like
> Alexa that would send the audio recorded on the device to NSA for
> processing.
> One of the problems is surely the complexity and hardware requirements to
> run such tool (I mean a dialog engine like Alexa).
> 
> However I do not understand why the linux community does not have a usable
> STT application as even 10+y ago there were usable tools for windows.
> And here we come to the original question. I have seen and tested some of
> the windows tools 10+y ago. They implement some type of learning, adapting
> themselves to the way you speak etc. In 2-3 months period the success rate
> goes to 95-98%. Price back then was affordable - AFAIR 200-400 US
> My advise, if you need such tool - look for commercial applications and
> forget the linux crap. I am willing to convince myself in the opposite, but
> looking what happens on the linux scene ... I've lost faith
> 

I abandoned proprietary software when Windows XP Pro SP3 was current.
Vendors think the user should only be interested in problems for which 
they have a $olution.

I have an absolute requirement that any software I acquire run under 
Linux - preferably Debian.

One focus of DeepSpeech seems to be aimed in my direction. It may not be 
as much of an "out of the box" solution as I would prefer. But it seems 
close enough to merit more investigation.

[toc] | [prev] | [next] | [standalone]


#228505 — Re: Feasibility of speech recognition for note taking on dedicated laptop?

FromDan Hitt <dan.hitt@gmail.com>
Date2020-11-06 17:40 +0100
SubjectRe: Feasibility of speech recognition for note taking on dedicated laptop?
Message-ID<B8cWl-4e8-1@gated-at.bofh.it>
In reply to#228496

[Multipart message — attachments visible in raw view] — view raw

On Thu, Nov 5, 2020 at 11:26 PM deloptes <deloptes@gmail.com> wrote:

> Richard Owlett wrote:
>
> .....
> > Are there now end-user, Debian compatible, dictation applications that
> > do NOT require proprietary software nor internet connectivity? My
> > internet searching turned up primarily old material or tool-set packages
> > packages aimed at programmers creating their own packages.
> >
> > Comments or suggestions?
>
> Again one of these topics, where people post about software they do not
> actually use.
>
> Let me comment here my impressions. I studied speech processing and wrote
> my
> thesis on dialog systems in 2007. Until about 2005 there were still some
> open source tools like ViaVoice by IBM. Basically all of this was dropped
> by 2010 - no idea why - might be something related to Google/Amazon, costs,
> patents or whatever else.
>
>
>
>
>
Hi Deloptes,

If your thesis is available online, could you please post a link to it?

TIA!

dan

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web