G2G: Discussion: What should be the policy for editing name variants? [closed]

+51 votes
1.8k views

Hi WikiTreers,

I’d like to start a discussion about what the policy should be for editing name variants in WikiTree’s name variant database. If you didn't hear, the name variants can now be edited on WikiTree directly.

Name variants are just alternative forms of a name (e.g., Katherine/Catherine, Bob/Robert). WikiTree uses them in Person Search, Find Matches, GEDCOMpare, and some other places. More detail is at Help:Name Variants.

Here are some questions to consider:

  1. Should we allow adding variants that don’t have any profiles yet? For example, should “Smth” be in the database as a variant of Smith even if nobody on WikiTree has the surname Smth? Or should we stick to variants that already have at least one person on the tree?
  2. Should variants include common translations or associations between names across languages? For example: John, Jean, Ivan, and Juan share the same root. Where do we draw the line between “variant” and “related name”?
  3. What about common misreadings – should those count as variants? For example, should a name with “n” have a variant with “u”, or only when there’s evidence they’re often confused?

If you have comments on any of the above, or additional ideas or questions, please post below.

- Jamie

in Policy and Style by Owl (926k points)

Could we get an update on the status of this policy discussion?

Hi L,

Here is the proposal for a new project – I'm working on a draft of the form.

https://www.wikitree.com/g2g/2014116/proposing-functional-project-manage-name-variants-database

27 Answers

+42 votes
Thanks for the opportunity to feedback on this.

My broad view is that we should be erring towards a narrower rather than wider field of variants. I'd like to illustrate my thoughts with three family surnames Wellock, Houseman and Booth which I often use to "test" anything from record sets to GDPR compliance.

The potential match lists can already be quite long and, as the number of profiles continues to increase so will the number of potential matches.

If a name is rarely mispelt, say Booth, it is, of course, possible to limit the search to an exact spelling, which is handy because there are currently 50 variants. The only variant i've seen in use is Boothman, and that only as a distinct surname, not a variant although clearly they have the same root.

But this doesn't work where a surname is commonly spelled in a variety of ways. Wellock is one of those surnames. Even as recently as the second half of the nineteenth century, I've seen examples of just one immediate family being listed as Wellock, Willock and Wallock in different records, so I need to search with variants. Unfortunately the matches list is then invariably dominated by Welch (there are 20,000 of them, compared to 350 Wellocks), a surname which I've not yet seen used in any records as an alternative. On the other hand, most of the alternatives I have seen in use in the records sit much further down the list, past the 10 to 15 surnames picked up in the matches process.

Houseman suffers similarly. The top variant is Cosman. I have no idea why. I have a feeling I tried to remove it from the werelate list. But Houseman is useful in relation to the second point you raise - translation or language variations. I can trace my own line (which dominates English Housemans) back four hundred years with suprisingly little variation in spelling. But in North America it's different. Here the modern Housemans ancestry is more often an anglicised version of Hausman (or similar) meaning those language variations are important. Although I do dread an increase in the number of John matches when adding in Jean etc.

Finally, I think it's worth considering how quickly duplicate lines might get picked up. WikiTree is getting stronger in terms of basic dates and locations. It's not that long ago that I created a duplicate Wellock profile, because the original had very different (incorrect) dates. But when I started to add his wife, I found her, and hence the duplicate him. So it was just one profile to merge away. So yes, too narrow a list of variants will miss an odd match, but too wide, and people will stop checking for any and perhaps be more at risk of creating a whole line of duplicates.

I hope this is useful

Natasha
by Owl (247k points)

+22 votes
Others will surely provide references to genealogical best practices around this question.  I can only respond from my own experience.

I have been searching (without success) for a quotation that was found written inside the cover of a Norfolk church register regarding different individuals' spellings of the various names common in that parish.  It was both amusing and an important reminder that until relatively recently the spelling of someone's name depended on the accent with which it was spoken and the ear of the person who was doing the writing.  A clerk who was born locally would write names completely differently from the curate who was educated far away.

Hopefully when a researcher finds multiple spellings of the same person's name across different official sources (birth records, marriage records, census returns, death records etc) this can be noted in the profile.  Many other platforms do not take into account known other names when searching for matches and WikiTree can do this more elegantly since it has the metadata easily accessible.

Any computational assistance that can be provided in the capture and collation of different spellings for a specific person would be valuable (e.g. when the alternative spellings are apparent in source citations could these be swept up into the "Other Names" field at some point in the automated profile creation or review process?)

Whether "Corden", "Cordan" and "Cordall" are all spellings for the same family's surname (as in the case of my ancestors) probably depends on the country, county and the point in history at which it was happening.

Hopefully the variant matching database will be smart enough to realise that Thomas Brown might be known as "Tom Brown" but Sarah Thomas would never be known as "Sarah Tom".

Finally many church registers before about 1800 were written in Latin so a person known as William could easily have been recorded in the baptismal register as "Gulielmus".
by G2G-9 (9.4k points)

'n Latin so a person known as William could easily have been recorded in the baptismal register as "Gulielmus"

and if mentioned as a parent may appear as "Gulielmi", in the genitive case. One confusing example was a Paul, not recorded as "Paulus" but as "Pauli" which was then misread as "Paull" and is now enshrined in some people's trees!


Mark great answer.

I wonder if an "other names" category might be added.  

I have Gilbert in my Cornwall tree.  It was transcribed as Jelbert. Anyone searching might find this family in the J listing and not know about the G listing.  It was a speaker and transcriber issue as you mentioned.

In a first name, "Hettie" was posted on a grave marker [FAG} for Hepsebah. Who would know this is the same person?

I'm glad there is a gene program like Lost Cousins where you can list the name as it's posted, and list the name as it was given. You are helping to correct the censuses. (censi??)

I've had people tell me G and J wouldn't have been interchangeable in medieval times because of the soft and hard sound.. please! I have so many examples of it.

+26 votes

Thanks Jamie, for raising the subject.

As member of some One Name studies, I have already listed variants of the surnames on the main Study page and in the variants worksheet of the One Name app.

So I used your link and went to check the variants of some surnames in this new variants database. Sadly I found just one name listed, which was incorrect (as far as I know) and none of the variants shown on the One Name Study page. Also, I could not change anything as this is limited to supervisors. Fair enough, but how do I find and contact a supervisor to make changes?

Apart from that you raise a very real concern: misreadings of, or miswritings in source texts. One of the names I am doing research on is Bezuidenhout. Although everybody now knows the correct spelling of this name (as shown on a map published in 1730), clerks in the 17th - 19th century did not and wrote down variants, like Besuidenhout, Bezuijdenhout, Bezuydenhout, Bezuidenhoudt and also any combination of these spelling variations. That being bad enough, transcribers also have tried to imitate handwriting by using the weird characters ú and ÿ , taken from a Unicode list, even in LNAB !!!

Please note that I am not against using such characters provided they exist in the alphabet used at that time, in that country, in print, not handwriting.

As a result, the number of variants of a name may explode, so I think it a good idea to agree on some rules.

  1. Only surnames already used as LNAB should be allowed as variants. 
  2. For first names, translations and shortened names should be allowed as variants, but not nicknames and modified names. So Lizzie or Betsie may be variants of Elizabeth, but not Albert who was baptised Allie but adapted a new, English name.
  3. Names may be incorrect written in source texts, or they may be written correct but transcribed incorrect. In both cases it should be checked if this error occurs repeatedly. My own surname Barnhoorn was often misspelled, but one variation persisted and resulted in a branch of the tree, with surname Barnhorn. So Barnhorn should be a variant, but Barrenhorn should not.
Hopefully this reply helps to forward this discussion. 

by G2G-10 (14.4k points)

There are languages where letters with accents simply belong to the languages. Let's take the Spanish name "Lopez". Outside of Spain the usual spelling is Lopez, but in Spanish it is López with the accent. Or in the Balkan you have names like "Bogunović". Yes, the Ć is a different letter than the usual C. So the accents are needed. You cannot simply "ban the accents". This does not work in a worldwide Tree.

+23 votes
As a starting point for last name variants, I would traverse the connected male line profiles and sibling profiles and create any variants based on what is found in the WikiTree database. Then take request(s) from the members as provided. I would limit the last name variants to what is found in the Wikitree database.

As for first name variants, I would use whatever sources we have today of what the variants have been found to be.
by Golden Owl (3.1m points)
edited by Tommy Buch

Please be very careful with this approach. It will only work where names are passed from the male line and not other systems are used. It will not work in a lot of non USA based places.

+21 votes
For 1, my instinct is to say yes, because obviously when someone with that variant shows up, we want to be sure it's not already on the tree in the other form.  but then again,  if it isn't on the tree already, why not?  make a profile with sources so there is an example to use in the system, and then request it as a variant.  

2) I think the "easy" ones obviously, but hmm, I'd have to think about how far out on translations to go--Michael and Mike and Mikael and Mischa are all pretty clearly the same name, and should.  And there are some very unclear but common variants like Mary and Polly that show up all the time in english language genealogy and confuse newts, and probably similar ones in other languages, like Sasha and Alexander.  But as we continue to get international, do we want 100000 Johns showing up when someone makes a Ivan?  probably not.

3. is near and dear to my heart, as my Guillou ancestors, especially the women, end up in trees as Guillon fairly often.  I doubt they were ever *called* this, but it's in marriage and death records.  i think for this sort of variant, we should ask for proof--something like records of at least two separate individuals with the mispelling variant.
by Owl (192k points)

+28 votes

My primary points of view on Surname variants hasn't changed much since we discussed this years ago Jaime. I would like to see wikitree apply the same drive to accuracy and required reasons for changes to the surname database as we expect from individual profiles. 

As main concepts:

1.  I think we should establish a baseline of variants for each surname. Wikitree has all the resources already available to understand which wikitreers are most active in which surnames. The new system should strive to utilize their expertise in creating that baseline. 

2. Once a baseline is established, some evidence MUST be provided to alter the database.

This is an example. Currently there are ~460 profiles with the surname FRATTA. Sysop/Leaders can easily see which wikitreers are most active creating/connecting those profiles. Some method of communication with those individuals for fact based input or giving those individuals some temporary powers to make corrections should be considered. 

In this example, the most active wikitreer would be me. I can categorically state that NONE of the 1000's of profiles supposedly related here have ever been found to have any connection whatsoever. The one variation ever found was the post-immigration change of one branch to FrattO.

Related Family Names (to FRATTA)

 For me the baseline would be Fratta and Fratto ONLY related to each other. Once that baseline was established, a SOURCE would be required for an alteration.

by Owl (113k points)

Hah Nick, I was thinking about contacting you soon since I've had a to-do item to contact you (on hold until the database was live) since June 2023!

I have a Leader interested in leading a Variant Names Project, so part of that project's tasks could be working with various ONS or interested individuals to clean up the variants. The project badge could be awarded temporarily to give editing powers.

Hi Nick,

I think your point 2) "2. Once a baseline is established, some evidence MUST be provided to alter the database." is crucial, but/and I wouldn't hold that evidence line too high, and there'd need to be a straightforward route to doing so.

So I like the idea of starting narrow and add to it.

I was interested in the list you shared, because Pratt is one of my historic surnames, and I've never seen it become a name starting with an F. It would be like a couple of the examples I shared as it would dominate.

So two things. I am not sure if alternates could be one directional? So an alternate to Pratt could be Fratta, but Fratta didn't come up with Pratt?

Two, if we started narrow, could we set up a seperate G2G request,similar to categories, although I am aware that doesn't have a seperate channel either at this point, and make it quite simple. Just share two sources with two variants on either one profile or two which were parent and child?

I am not sure if alternates could be one directional?

The variants are bi-directional. Adding or deleting a variant from a name will change the variant as well.


+18 votes

I would like to add another thank you for raising this. I am surprised I hadn't heard about the change before, having tried the WeRelate method some time ago and found it no longer worked. I have been frustrated by the inability to add variant names since I first started using WikiTree and found others creating duplicates of my profiles and occasionally the other way round, because of the lack of variants on the database. I get that it needs supervision, but it isn't clear who qualifies to add a variant. There are many examples of profiles with variants of my family name additional to the only one found on the variants database, but who can add them? The [[Help:Name Variants]] page says you need to be a Project Leader, but doesn't say what kind of Project. I have recently started a One Name Study/Project on WikiTree, but that doesn't seem to qualify me. The box at the top of the EditVariants page states that you need to be a sysop or supervisor to edit.  

All those responding to this question are raising issues that I have been of interest to me for some time (see: [[Space:Shakespeare Variant Names|Shakespeare Variant Names]]) and I realise that my concern is a little off theme here, but I would be grateful if someone could clarify who can do what.

With hope,

Martin Shakespeare

by G2G-4 (4.1k points)

Hi Martin,

Right now, the people with the Project Leader badge can edit the variants (although have been asked to be conservative about editing until we have some sort of policy in place). ONS badges don't count at the moment.

I do have a Leader willing to lead a Name Variants Project, so eventually a special "Name Variants Project" badge could give permissions for people other than Leaders to edit the database.

I think the research on your FSP is great! Would you recommend that all the variants in the "Shakespeare Variant Names on the Wikitree ONE NAME TREES App" section be added to the database?


Thanks for the response and your query about adding the names. Yes, I recommend that most be added, but I would like to go through the list first to double check that those variant names with only a few profiles can't be handled another way so that I avoid the need to have more variants than necessary added. I will message you when I have done this.

Thanks again.

Martin Shakespeare

+27 votes
I agree with those who want to keep the variants as narrow as possible.   As the lead of the Arborist project and as a mentor, I find that newer members when presented with a list of 30 or more potential matches see the review as a time consuming effort to go through all those matches.   For a new member it appears easier to create the profile and then eventually discover the duplicate.
by Owl (966k points)

+8 votes
Clarifying question Jaime: Is the name variant database only used for searching inside of Wikitree or is it also used by, For example, rootssearch to search external sites? My answer will be different depending on your answer.
by Golden Owl (1.1m points)

It's only used for WikiTree searches, and isn't used for searching other sites.

I am deeply grateful to everyone who is responding and to you folks who know a heck of a lot more than I do, who think deeply about these things.  I can barely follow parts of the arguments so I will sit back and gratefully use whatever is decided.

+21 votes
I confess to not having read all of the existing responses, so apologies if I'm repeating what someone has already said.

I think that as point zero, given names need to have different rules than surnames. Completely different: documentary practice in Europe, and to some degree in North America, treated the two in polar opposite ways.

Names have always been translated into the language being written. We still do it today: the English abstract for the Hungarian dissertation names its subject Pipo Ozorai, not Ozorai Pipo like the dissertation itself, and not Pipo de Ozora like the actual historical documents.

The exact details of this translation varied, but the version that's applicable to genealogy is where the given name is translated (or substituted, see below), while the surname is just along for the ride, modulo the orthographic variation caused by multilingual environments.

For example, in a German-language record, you find Anton Schneller; in a Hungarian one, he's Sneller Antal (or Schneller Antal, depending on the writer's knowledge, habits, and interpretation of the rules), and in a Latin record written in Croatia, he's Antonius Šneller. As you can see, for the given name, the precise sound is irrelevant, while for the surname, sound is all that matters.

Translation of given names sometimes disregards not only the precise sound, but also the etymology: Gyula becomes Julius, Béla becomes A(da)lbert, Rudolf becomes Rezső. Or Wojciech or Vojtěch become Adalbert. (Or add a -us to the German forms if it's Latin.) Should such substitutions be included in the variants database?

For surnames, I don't know how to address orthographic variation that depends on the language context for phonological interpretation, such as the directly-opposite use of 's' versus 'sz' in Polish versus Hungarian.

Ideally, I'd like to restrict surname variants to forms that are actually found in documents for the same person, or at least the same person's immediate relatives. There's no way to enforce that, though, neither in the existing lists nor for future suggestions.

I also don't know what to do about misreadings, such as Sueller instead of Šneller. Hopefully nobody's entering those in the name fields on WikiTree, right? They're not actually the person's name. They're useful to keep in mind when searching indexes of record repositories -- which WT is not. The question is complicated, though, because sometimes the misreading is a valid name, and it can take quite a bit of research to figure out what is actually meant by the consistently-ambiguous handwriting.
by Owl (177k points)

And yet sometimes those "misspellings" which could also have been "mis-hearings" become the later valid name.  We should not be guessing as to what the writer thought they heard compared to what they actually wrote.

In the 17th and 18th centuries this was true all over Acadie and Quebec for the french.

The question of mishearings is a whole 'nother kettle of fish. Or can of worms. I'm happy if I can figure out the handwriting-interpretation end of it. :-/

+18 votes

hi Jaime,

As long as we're using such a db, then for 

#1, include true variants in the db; I would not exclude such just because we don't have a profile on WT yet.  We're not finished.  wink

#2 Nonononooooooooooooooooo!  Translations are a big nono.  They happen when somebody migrates to another culture / language, but they should not be included in a db of variations.  They aren't a variation, they are a corresponding name in another language. 

#3 that one should not be a db item, ambiguous script when the letters N, U and V are easily confused for each other is a problem for paleographers.  (it's not just U and N, V is in there too)  One can only determine what the correct name is by looking at a number of documents in different hands very often.  U in French sometimes gets written ü (tréma on top, similar to German Umlaut) by some scribes, this is a deliberate attempt by them to remove the ambiguity of their handwriting.

by Owl (925k points)

This illustrates that we need to separate given names and surnames in these discussions, because translation of given names (#2) applies All The Time in places/times where records were in a different language than the vernacular, such as Latin in Roman Catholic registers.

nope, I beg to differ.  The priests wrote them in Latin, but the parents etc said them in their own language.  We don't use the Latin name in our tree, at least not around here.

I agree that the Latin (or German, or heck, English) name is not what should be entered on a profile if that's not the language that the parents used -- but this doesn't mean that people don't enter those names on profiles. Therefore, Person Search, Find Matches, and GEDCOMpare -- the parts of WikiTree that use the name variants -- need to be able to match up an existing Alexander with the Sándor I just entered, meaning that for given names, the database absolutely must contain translations. Otherwise, it's useless.

sorry, we're not going to agree on this one I don't think, to my mind having translations of given names in the database makes for horrendous results, as example Guillaume has the corresponding English name William, and I don't know how often I've had to restart a search because it would spew out all the translated names.

Reading this, and being particularly struck because I have a Guillaume/William ancestor, I'm thinking maybe this is a balance between a variant and what to use as alternative names in a profile.

Maybe a better alternative than adding first name variants is to better encourage people to add true alternate names where someone has been, say, born as a Jean in France and then adjusted to John in the US in later life?

Natasha, in cases where somebody migrates and is thereafter known by an anglicized version of their name, this would be covered by entering the proper given name at birth in the first box, and entering the later version in the preferred name field.  Or nickname field, depending on actual usage, of course.

That was what I was meaning, but failing to articulate!

What about Belgium, where someone's forename might appear in Dutch in one record and in French in another (plus Latin in a church record). Does the man who was christened Carolus Ludivicus, married as Karel Lodewijk, and was buried as Charles Louis need to have all those names in the name fields so that someone who has only one of those records can find him on WikiTree?

If I only know him as Charles Louis, how will I find him if he is only entered as Carolus Ludivicus or Karel Lodewijk?

yes, they would be there as far as French and Dutch.  Not the Latin version though.  Clergy used Latin, with variable time frame as to when they switched to the local ''vernacular''.  But the majority of folk did not speak Latin, so would not have called the person by a Laton version.

What matters to the algorithms or processes that use the variants database is not what names the people themselves used, but what names for them have been entered on WikiTree -- and trust me, that means the straight-from-the-Latin-register names (in their full genitive or ablative or other non-nominative glory) All The Time.

It's not unique to WikiTree that people don't recognize Latin names as Latin: for most people, one foreign language looks much the same as another, so they don't realize that their ancestor didn't change names between her baptism as Elisabeta and her marriage as Erzsébet.

What is unique to WT is an added layer: the "Last Name At Birth" label leads to a whole One True Name fallacy that's so well-enshrined that many (or possibly most!) users believe that it's WikiTree policy to enter a person's full name exactly as it appears on the earliest piece of paper about him. This means that the matching algorithms need to be able to find people as Elisabetha and Antonius, regardless of the fact that nobody ever called them that.

Can't speak for other languages, but French profiles don't follow what you are describing of using the Latin form of the name, at least in the large majority of cases I have seen.  But for given names, as opposed to last names, I don't think the db should include them.  Period.  Or else, have a separate small db for given name correspondences in multiple languages.

If the given names lists do not include multiple languages, we may as well stop running the automated checks for duplicates, because they'll only ever find one by sheer chance.

+16 votes
I would want a greater number of variants that would allow more matches.  The closest name matches appear at the top of the list when making a new profile or doing a general search.  More variants would allow for more matches.  The french names I deal with have many variants.

But I do agree with Robin, and so does choice theory (which we use in marketing).  Too many choices cause people to become tired and give up.  If we use a broader set, could we emphasize more the use of filters?  The filters are really helpful but they are tiny buttons that I hardly remember are there.  That should cut down on the realistic choices considerably.  Also IDK if the list provided on profile creation uses the data that might be entered there as to dates or locations, already.  If it did, that would also help provide a shorter list for review.
by Owl (547k points)

+15 votes

This seems to fall between two bounds: (1) too many variants make it difficult to extract the one item you need and (2) too little variants means you miss the "odd" one that you have been looking for years.

We already have the ability to switch off the variant search completely, so as the questioner proposes - how far should it go?

Much of this topic could easily slide into opinion, so thanks for the question, hopefully some scientific method will emerge.

First, to make the process as simple as possible I would only allow names that are already on wikitree - accepting that this might curtail the use of externally purchased/licenced lists - but it will stop people stuffing the lists just as a hobby in itself.

Second, only variants that appear in a "reliable source" should be permitted (ie BMD) - because that's the primary purpose of the search; to find a record that has been misspelt/misheard/mistrancribed. So for me, even if the variant is rare, it should be included as might be useful to someone else.

I am not keen on "special people only" being able to add/edit varients, however if you have to pick a group, the ONS coordinators would be the only logical  pick for last names as they usually have a better knowledge than most on their own topic. You would have to find a way to restrict edits to their own area. However, all changes are reversible, as with profiles; so let anyone edit the list, maybe limit the changes to one a week per user (or another period) to slow down crazy people and Spambots.

The best solution, in my opinion, would be to have more than one level of variant reflecting the "confidence" of the match; one odd case occurring could be "low" and several cases occurring could be "high". Wkitree's own data could ultimately be used to adjust the confidence level. However this would require changes to the search and how the data/indexes are held so that the confidence level could be set/tuned; this may be a bridge too far for wikitree management. That said, this would make WikiTree's search the best on the internet; no other site has this from my observations.

Just another thought on adding/editing variants - the system could permit/reject a change based on whether or not the variant appears in at least one hop; ie, if the parents or children of a profile have the variant it is permitted - this moves the editing back where it belongs; with the people that edit/manage the profiles.

by G2G-9 (9.3k points)

+14 votes

Some surnames have only a few very common variants.

My own name, Horsley/Horsely/Horseley, is one and names I follow include Acklam/Acklom and Lamplugh/Lamplough. It should be easy to incorporatthese variants in a search without adding major mismatches.

Forenames are a very different business.

David/Dave, is obvious but if you live in Wales or on the border it should be possible to add Daffyd/Taffy/Taff/Dai/Dewi without adding a slew of foreign translations.

Place names are also a problem; Kirby/Kirkby, Little Riston/Ruston Parva, Lampeter/Llanbedr.

Obvious, but not easy to implement.  Perhaps we could provide /a way for individuals to have a personal file of variants.

by Owl (38.9k points)

+14 votes
I think a surname variant can only be considered variant (rather than deviant) if there are multiple reliable primary source documents with the variant for a surname within one family group i.e. parents and children.
by Owl (117k points)

+13 votes

Seems to me that two or maybe three variants happen:

  1. Subject (i.e., person in a profile) or their descendant - with or without intention, changes the spelling for any number of reasons. For example, some Fann fathers had sons who used Fenn; Autry --> Autrey, etc.
  2. Some ancestors effect "changes" due to an inability to read or write/"knowing their letters," so the result may introduce dialect variation or recorder best practices for knowing what spelling was either phonetically intended or traditional - and the latter is often affected by rarity or familiarity with the subject's given dialect.
  3. Recorder or transcriber error, quite frequently due to wide variations in the quality of handwritten cursive.

I think source images (as opposed to summary transcriptions) become crucial to validate the value of some of the variants. Others just require a metareview or contextualizing all souces at hand.

Type 1 is typically proven by a shift in spelling for recorded sources between generations and is often reported in multigeneration family name genealogy or studies. Type 2 can sometimes be in evidence by frequent shifts in recorded spellings between sources for the same person. And, Type 3 often comes to light by noticing holes in the data, and doing wildcard searches that find spellings like Farm for Fann, or Daine for Bain.

The real question is the level of sensitivity that might be optimal for the name variant dbs: if the goal is to both avoid duplicates AND "standardize," then Type 3 "sources" might need to be flagged for review before being introduced into the dbs (i.e., flag the spelling as so rare that a layer of review is needed before adding the LNAB vs., the following similar spellings exist and warrant more research before adding a LNAB variant: please submit your sources with images preferred, to justify this different LNAB addition).

The first two types of variation might require less scrutiny, so long as the user is alerted that alternate LNAB instances are known, and perhaps the user can tick a box that they examined more than one generation of the surname variant/express a level of confidence or choose whether they have evidence or an indication of Type 1 or Type 2 accounting for the variation.

It's difficult to balance the demands for managing this policy. Sure, virtually everyone accepts one world tree, and is therefore commited to avoiding duplicates. Note, however, that access to the Other Names field to document variants seen in usage is not possible until AFTER the profile is created. Having a dedicated project leader, and all leaders able to approve additions might be sufficient. 

Asking users to provide image sources before adding a variant person's profile might seem onerous, but everyone stands to benefit from a multigeneration perspective and review; however, it may dissuade some users who haven't encountered this level of analysis, and might be especially frustrating for those used to relying on standard FamilySearch transcriptions and who never actually examine images, which are not always readily attached to a source link.

by Owl (128k points)

as far as I know we do not ''standardize'', since those supposed ''standards'' change depending on who is deciding about them.  For instance, Between 2 sources, Tanguay's dictionnary and PRDH database, there are names that were ''standardized'' by either one in different ways.  Who's right?  Frankly, neither, if one looks at the actual records.  For instance, the name Gauthier (PRDH's ''standardized'' name) is seen with various spellings:  Gaultier, Gautier, Gauthier, Gonthier, Gotier....  and changes over time.  

I've seen ''standardized'' names that were utterly misleading.  Such as one the filles à marier, whose name was written Macré, Maqueray..., PRDH changes the name to Macray.  She was French, not Scottish, so this is totally misleading.  Finding records in France with the name Macray would be impossible.  We do NOT use such.

Yeah, no, I didn't mean standardize in the sense that you describe, which is why I put "standardize," meaning "test whether the presently submitted, new LNAB variant is a true LNAB or require further source review occurs and images need examination before allowing the system to automatically allow the new variant to be entered, without further justification by the end user."

Poor word choice, but I couldn't think of a better single word choice to imply all of the above. Maybe "novel variant testing?"

It becomes a choice.  That a name got spelled one way at the baptism of a child (which is where most LNAB originate), would get entered in the appropriate box.  If a person then uses another variant of the name, that gets entered in either CLN or OLN boxes.  I don'T see that the variant database should be a criterion for such.  It's a tool for searches and GEDCom upload comparisons.  Not for determining what is correct for individual profiles.

Despite its ever-growing size, WikiTree is entirely missing more surnames than it contains, by several orders of magnitude. Impedance on entering profiles with new-to-WT surnames would be ill-advised, and the prior existence -- or lack thereof -- of profiles with a name should not be considered relevant to the database of name variants.

+15 votes
Jamie,

Without having read all the answers that are already have been given in this topic, the first question that pops up in my head is… Why are we doing this? Why can’t we use a wildcard like * for searching in names like almost all other use (at least in the Netherlands). Then you won’t need this database at all.

My other remark is that if only project leaders are allowed to add names to this table, who (which project lead) should you need to contact for example for a Dutch, Notable, Holocaust person that is also part of a one name study?
by Owl (286k points)

You know, that's a good question: why _doesn't_ WT use wildcards in user-input searches?

But I cannot think of a way to apply wildcards to non-user-input searches, such as WT's automated checks for existing profiles at new profile creation.

Well my idea was related to the search functionality, not to the creation of new profiles.

You can use the "*" or "?" wildcards in searches, and have been able to for several years.

We've also been using variants for as long as I can recall. It's just easier for the variants to be edited now, instead of having people go to another site to edit them.

There is a plan for there to be a project that will maintain the list. So you would contact that project.

Wowww, that’s fantastic!! Why didn’t anyone told me this before? What’s the difference between * and ? In the search at WT?

"*" is zero or more characters, and "?" is one character.

So N*son would match "Nson", "Nelson" or "Nabcdefgson", etc. "N?son" would match "Nison", "Nason", "Nuson", but not any of the examples for "N*son".

The name variants lists are relevant to new profile creation, as the automated check (that I mentioned) uses the variants, and (as I said) I can't think of a way for that process to switch to wildcards.

But it's news to me, too, that WT does know wildcards, and specifically the asterisk and question mark (not the SQL-based percent and underscore that for example the Hungarian genealogy society's site uses). When was this capability added, and why isn't it mentioned anywhere that I can find?

Thank you very much Jamie, I’m happy that I now know how to be able to search with those word cards. Much appreciated

Hi J. Palotay,

I understand now more the use of those variants tables. Thank you.

I looked it up and it was in 2018 (I knew it was around the time I joined the Tech Team). There is a section on Help:Search: https://www.wikitree.com/wiki/Help:Search#Wildcards

+12 votes
New Netherland profiles are a particular puzzle.  We have people of lots of different origins being recorded by Dutch, German, English, and French speaking clerics, using patronymics, family names, and occasional occupational names.  For an extreme example see, "Harmen Meyndertsz (van den Bogaert) van der Bogart (abt.1612-abt.1648),"  (https://www.wikitree.com/wiki/Van_den_Bogaert-20 ).
by Golden Owl (1.5m points)

Same in the early years of Cape of Good Hope (now South Africa). Most people who arrived there were first registered with patronymic and often also place of origin. As a result very many souls are entered in wikitree with a patronymic for LNAB, even when they were later known with a surname. Surely the patronymic must not be a variant of the surname!

Also, for a long time, very many spelling variations appear by mispronouncing, mishearing, miswriting, misreading or incorrect transcriptions of handwriting. Surprisingly not so much in First Names, but a lot in Last Names,  My solution for improving searches would not be to add all these variants to a database, but to use the properly spelled name as LNAB and put these variants in Other Last Names. (as is already done in some profiles).


sorry Dick, but I have to disagree, the records that use a phonetic rendering of a name are often due to change of culture / language.  English or other origin migrants to New France for example get their names changed considerably.  The children don't get the original name but the phonetic variant heard by clergy, so it would be false to assign them the original LNAB.  These names are often found in subsequent generations, you basically have a new family name created this way.  As I've pointed out elsewhere, names also evolve over time.

Thanks Danielle, I am glad you agree with me that spelling variations that persist over one or more generations should be treated as a new family name. At the same time they belong to the same family and may be considered to be variants of the original name.

My comment was inspired by people having a different spelling (not a translation) for their name by every writer involved (and by transcribers introducing new characters to mimic handwriting). So when you go searching for a person, you don't know which of the 25 odd spellings to use. Fortunately I just learned in this G2G item that wildcards can be used (didn't know that). For one name I am researching (Bezuidenhout) none of the misspellings persisted. The original spelling is now used by everyone of that name. For other names there are branches (Barnhoorn=>Barnhorn, Blois=>Bloois, Stegenga=>Steginga)

+11 votes
  1. Variants across languages should be considered, since this is common at least in the U.S. when immigrants (or their first generation American children) sometimes anglicize their given name.
  2. Also to consider is Ireland's Roman Catholic case where a Latin-ish name is used in church records; a different common variation used in "real life"; and different common name closer to the Latin name gets adopted from the Church record into the Civil record. Depending on which record(s) the WikiTree entry is based on, any of these names may end up on the profile.
    e.g. Eugenius, Hugh (or Owen), and Eugene.
As for adding variants, I assume I should make a request to a Project Leader such as the Irish Roots Project for something like #2 above?
by Owl (190k points)

The association of etymologically-unrelated names (such as Hugh and Owen with Eugene) occurs in every language, as does the dissociation of related names (such as James and Jacob or Isabella and Elisabeth).

+11 votes

I think we need to keep the Name Variants system manageable and take a balanced approach that keeps it useful without making it overly broad. To do that, I suggest we treat Given Names and Surnames separately and establish clearer criteria for both.

Separate Rules for Given Names and Surnames

Given names and surnames function differently in historical records. Because of that, they should not be governed by the same standard. Creating distinct guidelines for each would help avoid confusion and conflict.

Surnames

For surnames, I recommend a narrow and evidence-based standard:

  • Variants should be limited to names actually recorded for that individual or family in WikiTree profiles and supported by reliable sources.
  • Variants should reflect historically documented spellings within the same lineage.
  • Misreadings, transcription errors, or clear recording mistakes should not be added as variants.

To support this, we should define what “evidence-based” means. For example:

  • The spelling appears in primary records connected to the individual or family.
  • The spelling appears consistently in credible secondary sources.
  • The spelling is historically documented as a variant within that lineage.

Each added variant should include a brief note or citation explaining the evidence. That way, future reviewers understand why it was included.

Given Names

For given names, I suggest keeping the scope limited to spelling variations of the same name.

  • Variants should generally reflect spelling differences of the same name. Like James/Jim
  • Cross-language equivalents should not automatically be treated as variants.
Because WikiTree is an international community, treating cross-language equivalents as variants can quickly lead to cultural and linguistic disputes. Keeping the focus on spelling variation rather than translation may reduce conflict and keep the system more consistent.

Governance and Transparency

I also think we should review how variants are managed.

Limiting editing authority solely to Project Leaders may be too narrow and can reduce transparency. At a minimum, we should provide clearer instructions on:

  • How members can request changes
  • What criteria are used for approval
  • How decisions are documented

by Owl (273k points)

Related questions

...