Index
2017-04-12 09:55Ken Dibble : Fuzzy Name Searching
2017-04-12 10:33Alan Bourke : Re: Fuzzy Name Searching
2017-04-12 11:22Peter Cushing : Re: Fuzzy Name Searching
2017-04-12 11:49Stephen Russell : Re: Fuzzy Name Searching
2017-04-12 14:31Garrett Fitzgerald : Re: Fuzzy Name Searching
2017-04-12 15:18Ted Roche : Re: Fuzzy Name Searching
2017-04-12 15:41Garrett Fitzgerald : Re: Fuzzy Name Searching
2017-04-12 15:43Gene Wirchenko : Re: Fuzzy Name Searching
2017-04-12 16:22Mike : Re: Fuzzy Name Searching
2017-04-13 10:05Ken Dibble : Re: Fuzzy Name Searching
Back to top
Fuzzy Name Searching

Author: Ken Dibble

Posted: 2017-04-12 09:55:27   Link

Hi folks,

I've been thinking of how I can improve the ability of my users to

find people's names in a system that has over 30,000 people in it.

I've looked at soundex, and I've considered munging names to remove

spaces, apostrophes, hyphens, etc. The thing about those approaches

is that in order to be efficient, they require pre-processing all of

the names in the system and storing the results, which can then be

queried to find matches.

Unfortunately, that would require modifications to the database,

which I try to avoid due to the downtime they require.

I'm looking for suggestions on how to produce results that include

close matches on last names that doesn't require pre-processing.

I've played with various schemes to assign "weights" to matches based

on the number of matching letters, but they all end up being very

slooooow and also producing too many false positives.

I suppose there are no easy answers, but if anyone has an algorithm

for this kind of thing that they would be willing to share, I'd be grateful.

Thanks.

Ken Dibble

www.stic-cil.org

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/31.FF.16480.FDF3EE85@cdptpa-omsmta02

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Ken Dibble
Back to top
Re: Fuzzy Name Searching

Author: Alan Bourke

Posted: 2017-04-12 10:33:31   Link

FoxWeb has a full text search engine, free, that might help.

http://www.foxweb.com/fwFullText/

--

Alan Bourke

alanpbourke (at) fastmail (dot) fm

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/1492011211.3756611.942605856.1911D0E1@webmail.messagingengine.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Alan Bourke
Back to top
Re: Fuzzy Name Searching

Author: Peter Cushing

Posted: 2017-04-12 11:22:47   Link

On 12/04/2017 15:55, Ken Dibble wrote:

> <snip>

> I'm looking for suggestions on how to produce results that include

> close matches on last names that doesn't require pre-processing.

>

About 20 years ago I did some work on marketing databases and one of the

big tasks is de-duping data. We did things like substituting vowels

with a * in the search value, then you get a more general match that

avoids some spelling mistakes. Unfortunately approaches like this would

require you to store the processed data.

Do you have the surname is a separate field with an index? If so you

could show the results of matches in a box below the search which

changes as they type each letter. The user can then see the list going

smaller as they type and will hopefully see what they are looking for.

If they have typed the whole name in and can't see the match they can

remove letters to look for a partial match. I would only match from say

the 3rd character or more as it may be too slow to display all the matches.

Peter

This communication is intended for the person or organisation to whom it is addressed. The contents are confidential and may be protected in law. Unauthorised use, copying or disclosure of any of it may be unlawful. If you have received this message in error, please notify us immediately by telephone or email.

www.whisperingsmith.com

Whispering Smith Ltd Head Office:61 Great Ducie Street, Manchester M3 1RR.

Tel:0161 831 3700

Fax:0161 831 3715

London Office:17-19 Foley Street, London W1W 6DW Tel:0207 299 7960

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/19146d05-62bb-fe7d-33b4-2d11ffecccf3@whisperingsmith.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Peter Cushing
Back to top
Re: Fuzzy Name Searching

Author: Stephen Russell

Posted: 2017-04-12 11:49:16   Link

I remember this joy of searching names in a system that had 2+ million

customers and names were all varchar() instead of a key to a secondary

table. My indexes sure took a beating when I got another "Williams", the

number one last name in the system, and it had to tear a page to make a new

page in this area.

I found that making a table called NAMES fixed the search time I was

experiencing. Two text boxes had input for whatever they keyed. I added

the % for wildcard after any text in each box and one of the keypress

events was the trigger to run it.

Select <field_list>

from customer

where lNameID in (

select nameID from names

where Name like @Lname)

and

fNameID in (

select nameID from names na

where na.Name like @Fname)

That has been 10-13 years ago.

On Wed, Apr 12, 2017 at 9:55 AM, Ken Dibble <krdibble@stny.rr.com> wrote:

> Hi folks,

>

> I've been thinking of how I can improve the ability of my users to find

> people's names in a system that has over 30,000 people in it.

>

> I've looked at soundex, and I've considered munging names to remove

> spaces, apostrophes, hyphens, etc. The thing about those approaches is that

> in order to be efficient, they require pre-processing all of the names in

> the system and storing the results, which can then be queried to find

> matches.

>

> Unfortunately, that would require modifications to the database, which I

> try to avoid due to the downtime they require.

>

> I'm looking for suggestions on how to produce results that include close

> matches on last names that doesn't require pre-processing.

>

> I've played with various schemes to assign "weights" to matches based on

> the number of matching letters, but they all end up being very slooooow and

> also producing too many false positives.

>

> I suppose there are no easy answers, but if anyone has an algorithm for

> this kind of thing that they would be willing to share, I'd be grateful.

>

> Thanks.

>

> Ken Dibble

> www.stic-cil.org

>

>

[excessive quoting removed by server]

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/CAJidMYJnLyivczk6edHPTfbjcU_emJysfGKNy071aK6aiaC73w@mail.gmail.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Stephen Russell
Back to top
Re: Fuzzy Name Searching

Author: Garrett Fitzgerald

Posted: 2017-04-12 14:31:08   Link

I wrote a FLL to do Levenshtein distances for fuzzy name matching, but

everything was posted to my blog, which is no longer online. It wasn't

amazingly hard to figure out, though, so it might be worth finding the

algorithm in C and recreating my steps. It ran much faster than equivalent

Fox code did.

On Wed, Apr 12, 2017 at 12:49 PM, Stephen Russell <srussell705@gmail.com>

wrote:

> I remember this joy of searching names in a system that had 2+ million

> customers and names were all varchar() instead of a key to a secondary

> table. My indexes sure took a beating when I got another "Williams", the

> number one last name in the system, and it had to tear a page to make a new

> page in this area.

>

> I found that making a table called NAMES fixed the search time I was

> experiencing. Two text boxes had input for whatever they keyed. I added

> the % for wildcard after any text in each box and one of the keypress

> events was the trigger to run it.

>

> Select <field_list>

> from customer

> where lNameID in (

> select nameID from names

> where Name like @Lname)

> and

> fNameID in (

> select nameID from names na

> where na.Name like @Fname)

>

> That has been 10-13 years ago.

>

>

>

>

> On Wed, Apr 12, 2017 at 9:55 AM, Ken Dibble <krdibble@stny.rr.com> wrote:

>

> > Hi folks,

> >

> > I've been thinking of how I can improve the ability of my users to find

> > people's names in a system that has over 30,000 people in it.

> >

> > I've looked at soundex, and I've considered munging names to remove

> > spaces, apostrophes, hyphens, etc. The thing about those approaches is

> that

> > in order to be efficient, they require pre-processing all of the names in

> > the system and storing the results, which can then be queried to find

> > matches.

> >

> > Unfortunately, that would require modifications to the database, which I

> > try to avoid due to the downtime they require.

> >

> > I'm looking for suggestions on how to produce results that include close

> > matches on last names that doesn't require pre-processing.

> >

> > I've played with various schemes to assign "weights" to matches based on

> > the number of matching letters, but they all end up being very slooooow

> and

> > also producing too many false positives.

> >

> > I suppose there are no easy answers, but if anyone has an algorithm for

> > this kind of thing that they would be willing to share, I'd be grateful.

> >

> > Thanks.

> >

> > Ken Dibble

> > www.stic-cil.org

> >

> >

[excessive quoting removed by server]

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/CAGd8Mrc8sFvM=4Q4EoAu=PnJgyFCseByt2y_0ZPUJC2tB2=TXw@mail.gmail.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Garrett Fitzgerald
Back to top
Re: Fuzzy Name Searching

Author: Ted Roche

Posted: 2017-04-12 15:18:44   Link

Ah! The algorithm rang a bell!

Garrett: have you tried searching Archive.org? A LOT of your stuff

appears archived:

https://web-beta.archive.org/web/*/garrett%20fitzgerald%20

The equivalent FoxPro code, by the way is in the leafe downloads at

https://leafe.com/dls/vfp. Bob Calco wrote it up.

Also on Fox Wikis at: http://fox.wikis.com/wc.dll?Wiki~LevenshteinAlgorithm

Craig Boyd's blog about Spell Checking at

http://www.sweetpotatosoftware.com/spsblog/CommentView.aspx?guid=8800bdb9-a9c2-484f-942f-6a08947d903a

On Wed, Apr 12, 2017 at 3:31 PM, Garrett Fitzgerald

<sarekofvulcan@gmail.com> wrote:

> I wrote a FLL to do Levenshtein distances for fuzzy name matching, but

> everything was posted to my blog, which is no longer online. It wasn't

> amazingly hard to figure out, though, so it might be worth finding the

> algorithm in C and recreating my steps. It ran much faster than equivalent

> Fox code did.

>

> On Wed, Apr 12, 2017 at 12:49 PM, Stephen Russell <srussell705@gmail.com>

> wrote:

>

>> I remember this joy of searching names in a system that had 2+ million

>> customers and names were all varchar() instead of a key to a secondary

>> table. My indexes sure took a beating when I got another "Williams", the

>> number one last name in the system, and it had to tear a page to make a new

>> page in this area.

>>

>> I found that making a table called NAMES fixed the search time I was

>> experiencing. Two text boxes had input for whatever they keyed. I added

>> the % for wildcard after any text in each box and one of the keypress

>> events was the trigger to run it.

>>

>> Select <field_list>

>> from customer

>> where lNameID in (

>> select nameID from names

>> where Name like @Lname)

>> and

>> fNameID in (

>> select nameID from names na

>> where na.Name like @Fname)

>>

>> That has been 10-13 years ago.

>>

>>

>>

>>

>> On Wed, Apr 12, 2017 at 9:55 AM, Ken Dibble <krdibble@stny.rr.com> wrote:

>>

>> > Hi folks,

>> >

>> > I've been thinking of how I can improve the ability of my users to find

>> > people's names in a system that has over 30,000 people in it.

>> >

>> > I've looked at soundex, and I've considered munging names to remove

>> > spaces, apostrophes, hyphens, etc. The thing about those approaches is

>> that

>> > in order to be efficient, they require pre-processing all of the names in

>> > the system and storing the results, which can then be queried to find

>> > matches.

>> >

>> > Unfortunately, that would require modifications to the database, which I

>> > try to avoid due to the downtime they require.

>> >

>> > I'm looking for suggestions on how to produce results that include close

>> > matches on last names that doesn't require pre-processing.

>> >

>> > I've played with various schemes to assign "weights" to matches based on

>> > the number of matching letters, but they all end up being very slooooow

>> and

>> > also producing too many false positives.

>> >

>> > I suppose there are no easy answers, but if anyone has an algorithm for

>> > this kind of thing that they would be willing to share, I'd be grateful.

>> >

>> > Thanks.

>> >

>> > Ken Dibble

>> > www.stic-cil.org

>> >

>> >

[excessive quoting removed by server]

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/CACW6n4v99K4zgi0MTOD6Ng2adfbroy4XX-_Em90K=peF9O_DQw@mail.gmail.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Ted Roche
Back to top
Re: Fuzzy Name Searching

Author: Garrett Fitzgerald

Posted: 2017-04-12 15:41:23   Link

Aha in turn!

https://web-beta.archive.org/web/20101216062156/http://blog.donnael.com:80/2008/04/creating-an-fll/#more-1965

On Apr 12, 2017 4:19 PM, "Ted Roche" <tedroche@gmail.com> wrote:

> Ah! The algorithm rang a bell!

>

> Garrett: have you tried searching Archive.org? A LOT of your stuff

> appears archived:

> https://web-beta.archive.org/web/*/garrett%20fitzgerald%20

>

> The equivalent FoxPro code, by the way is in the leafe downloads at

> https://leafe.com/dls/vfp. Bob Calco wrote it up.

>

> Also on Fox Wikis at: http://fox.wikis.com/wc.dll?

> Wiki~LevenshteinAlgorithm

>

> Craig Boyd's blog about Spell Checking at

> http://www.sweetpotatosoftware.com/spsblog/CommentView.aspx?guid=

> 8800bdb9-a9c2-484f-942f-6a08947d903a

>

> On Wed, Apr 12, 2017 at 3:31 PM, Garrett Fitzgerald

> <sarekofvulcan@gmail.com> wrote:

> > I wrote a FLL to do Levenshtein distances for fuzzy name matching, but

> > everything was posted to my blog, which is no longer online. It wasn't

> > amazingly hard to figure out, though, so it might be worth finding the

> > algorithm in C and recreating my steps. It ran much faster than

> equivalent

> > Fox code did.

> >

> > On Wed, Apr 12, 2017 at 12:49 PM, Stephen Russell <srussell705@gmail.com

> >

> > wrote:

> >

> >> I remember this joy of searching names in a system that had 2+ million

> >> customers and names were all varchar() instead of a key to a secondary

> >> table. My indexes sure took a beating when I got another "Williams",

> the

> >> number one last name in the system, and it had to tear a page to make a

> new

> >> page in this area.

> >>

> >> I found that making a table called NAMES fixed the search time I was

> >> experiencing. Two text boxes had input for whatever they keyed. I

> added

> >> the % for wildcard after any text in each box and one of the keypress

> >> events was the trigger to run it.

> >>

> >> Select <field_list>

> >> from customer

> >> where lNameID in (

> >> select nameID from names

> >> where Name like @Lname)

> >> and

> >> fNameID in (

> >> select nameID from names na

> >> where na.Name like @Fname)

> >>

> >> That has been 10-13 years ago.

> >>

> >>

> >>

> >>

> >> On Wed, Apr 12, 2017 at 9:55 AM, Ken Dibble <krdibble@stny.rr.com>

> wrote:

> >>

> >> > Hi folks,

> >> >

> >> > I've been thinking of how I can improve the ability of my users to

> find

> >> > people's names in a system that has over 30,000 people in it.

> >> >

> >> > I've looked at soundex, and I've considered munging names to remove

> >> > spaces, apostrophes, hyphens, etc. The thing about those approaches is

> >> that

> >> > in order to be efficient, they require pre-processing all of the

> names in

> >> > the system and storing the results, which can then be queried to find

> >> > matches.

> >> >

> >> > Unfortunately, that would require modifications to the database,

> which I

> >> > try to avoid due to the downtime they require.

> >> >

> >> > I'm looking for suggestions on how to produce results that include

> close

> >> > matches on last names that doesn't require pre-processing.

> >> >

> >> > I've played with various schemes to assign "weights" to matches based

> on

> >> > the number of matching letters, but they all end up being very

> slooooow

> >> and

> >> > also producing too many false positives.

> >> >

> >> > I suppose there are no easy answers, but if anyone has an algorithm

> for

> >> > this kind of thing that they would be willing to share, I'd be

> grateful.

> >> >

> >> > Thanks.

> >> >

> >> > Ken Dibble

> >> > www.stic-cil.org

> >> >

> >> >

[excessive quoting removed by server]

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/CAGd8MrdxDACMRJgA-PG9Z=X-oOHBAqYBVzdZ-KO8jn5giwqqBw@mail.gmail.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Garrett Fitzgerald
Back to top
Re: Fuzzy Name Searching

Author: Gene Wirchenko

Posted: 2017-04-12 15:43:04   Link

At 07:55 2017-04-12, Ken Dibble <krdibble@stny.rr.com> wrote:

>Hi folks,

>

>I've been thinking of how I can improve the ability of my users to

>find people's names in a system that has over 30,000 people in it.

>

>I've looked at soundex, and I've considered munging names to remove

>spaces, apostrophes, hyphens, etc. The thing about those approaches

>is that in order to be efficient, they require pre-processing all of

>the names in the system and storing the results, which can then be

>queried to find matches.

>

>Unfortunately, that would require modifications to the database,

>which I try to avoid due to the downtime they require.

Why would that be an issue of consequence?

You add some columns to a table. The rest of the software can

ignore them. (Unless you use select * or other black arts, said rest

might never see the new columns.)

You can split up the task.

Write your code for filling in the new columns in your

add/change code. Then, write a utility to fill in the rest. Then,

implement the searching.

>I'm looking for suggestions on how to produce results that include

>close matches on last names that doesn't require pre-processing.

I can not see that the preprocessing would be very involved.

>I've played with various schemes to assign "weights" to matches

>based on the number of matching letters, but they all end up being

>very slooooow and also producing too many false positives.

>

>I suppose there are no easy answers, but if anyone has an algorithm

>for this kind of thing that they would be willing to share, I'd be grateful.

There are not, because different languages assign different

values to the Roman alphabet characters. You are going to have

decide on language trade-offs.

Sincerely,

Gene Wirchenko

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/2eb6450d3c82b371fab23887f16e244e@mtlp000086

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Gene Wirchenko
Back to top
Re: Fuzzy Name Searching

Author: Mike

Posted: 2017-04-12 16:22:48   Link

I found a Levenshtein function somewhere last year and have been using

it with MariaDB as a function on the MariaDB server, called from my VFP

9 application. It's exceptionally fast and works pretty well.

My application needs to get "as close as" matches to a random string

(for manufacturer product SKUs, which can be any length, any

alphanumeric mishmash...example: RGB745WEHWW).

Sometimes the user enters everything right except one character, and

this function returns a weighted list of "as close as I can find" known

SKUs from a table of 45,000+.

I can send the function to anyone who is interested.

Here's an example of how I call it (the Levenshtien function is named

klose, pcSearch is the text string to search for)

select sku, klose(sku,?pcSearch) as score

from (select sku from skus where soundex(sku) like

soundex(?pcSearch)) as hits

order by score desc limit 10

Mike Copeland

Garrett Fitzgerald wrote:

> I wrote a FLL to do Levenshtein distances for fuzzy name matching, but

> everything was posted to my blog, which is no longer online. It wasn't

> amazingly hard to figure out, though, so it might be worth finding the

> algorithm in C and recreating my steps. It ran much faster than equivalent

> Fox code did.

>

> On Wed, Apr 12, 2017 at 12:49 PM, Stephen Russell <srussell705@gmail.com>

> wrote:

>

>> I remember this joy of searching names in a system that had 2+ million

>> customers and names were all varchar() instead of a key to a secondary

>> table. My indexes sure took a beating when I got another "Williams", the

>> number one last name in the system, and it had to tear a page to make a new

>> page in this area.

>>

>> I found that making a table called NAMES fixed the search time I was

>> experiencing. Two text boxes had input for whatever they keyed. I added

>> the % for wildcard after any text in each box and one of the keypress

>> events was the trigger to run it.

>>

>> Select <field_list>

>> from customer

>> where lNameID in (

>> select nameID from names

>> where Name like @Lname)

>> and

>> fNameID in (

>> select nameID from names na

>> where na.Name like @Fname)

>>

>> That has been 10-13 years ago.

>>

>>

>>

>>

>> On Wed, Apr 12, 2017 at 9:55 AM, Ken Dibble <krdibble@stny.rr.com> wrote:

>>

>>> Hi folks,

>>>

>>> I've been thinking of how I can improve the ability of my users to find

>>> people's names in a system that has over 30,000 people in it.

>>>

>>> I've looked at soundex, and I've considered munging names to remove

>>> spaces, apostrophes, hyphens, etc. The thing about those approaches is

>> that

>>> in order to be efficient, they require pre-processing all of the names in

>>> the system and storing the results, which can then be queried to find

>>> matches.

>>>

>>> Unfortunately, that would require modifications to the database, which I

>>> try to avoid due to the downtime they require.

>>>

>>> I'm looking for suggestions on how to produce results that include close

>>> matches on last names that doesn't require pre-processing.

>>>

>>> I've played with various schemes to assign "weights" to matches based on

>>> the number of matching letters, but they all end up being very slooooow

>> and

>>> also producing too many false positives.

>>>

>>> I suppose there are no easy answers, but if anyone has an algorithm for

>>> this kind of thing that they would be willing to share, I'd be grateful.

>>>

>>> Thanks.

>>>

>>> Ken Dibble

>>> www.stic-cil.org

>>>

>>>

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/fb983af6-5c07-9184-ca20-96d81a0dd37a@ggisoft.com

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Mike
Back to top
Re: Fuzzy Name Searching

Author: Ken Dibble

Posted: 2017-04-13 10:05:23   Link

Thank you everybody. I will be working through these suggestions and

let you know what I come up with.

Ken

>I remember this joy of searching names in a system that had 2+ million

>customers and names were all varchar() instead of a key to a secondary

>table. My indexes sure took a beating when I got another "Williams", the

>number one last name in the system, and it had to tear a page to make a new

>page in this area.

>

>I found that making a table called NAMES fixed the search time I was

>experiencing. Two text boxes had input for whatever they keyed. I added

>the % for wildcard after any text in each box and one of the keypress

>events was the trigger to run it.

>

>Select <field_list>

>from customer

>where lNameID in (

>select nameID from names

>where Name like @Lname)

>and

>fNameID in (

>select nameID from names na

>where na.Name like @Fname)

>

>That has been 10-13 years ago.

>

>

>

>

>On Wed, Apr 12, 2017 at 9:55 AM, Ken Dibble <krdibble@stny.rr.com> wrote:

>

> > Hi folks,

> >

> > I've been thinking of how I can improve the ability of my users to find

> > people's names in a system that has over 30,000 people in it.

> >

> > I've looked at soundex, and I've considered munging names to remove

> > spaces, apostrophes, hyphens, etc. The thing about those approaches is that

> > in order to be efficient, they require pre-processing all of the names in

> > the system and storing the results, which can then be queried to find

> > matches.

> >

> > Unfortunately, that would require modifications to the database, which I

> > try to avoid due to the downtime they require.

> >

> > I'm looking for suggestions on how to produce results that include close

> > matches on last names that doesn't require pre-processing.

> >

> > I've played with various schemes to assign "weights" to matches based on

> > the number of matching letters, but they all end up being very slooooow and

> > also producing too many false positives.

> >

> > I suppose there are no easy answers, but if anyone has an algorithm for

> > this kind of thing that they would be willing to share, I'd be grateful.

> >

> > Thanks.

> >

> > Ken Dibble

> > www.stic-cil.org

> >

> >

[excessive quoting removed by server]

_______________________________________________

Post Messages to: ProFox@leafe.com

Subscription Maintenance: http://mail.leafe.com/mailman/listinfo/profox

OT-free version of this list: http://mail.leafe.com/mailman/listinfo/profoxtech

Searchable Archive: http://leafe.com/archives/search/profox

This message: http://leafe.com/archives/byMID/profox/E2.1F.03423.4B39FE85@cdptpa-omsmta01

** All postings, unless explicitly stated otherwise, are the opinions of the author, and do not constitute legal or medical advice. This statement is added to the messages for those lawyers who are too stupid to see the obvious.

©2017 Ken Dibble