Text
Text helpers for Brazilian names, company names and addresses.
Capitalize
- JavaScript library
- Python library
- Go library59 cases fail
- Ruby library79 cases fail
- Rust library
- .NET library79 cases fail
- Erlang library
Capitalizes the first letter of each word the way a Brazilian name, company name or address is written.
For example, "jose da silva" becomes "Jose da Silva", "empresa ltda" becomes "Empresa LTDA" and "santana/rs" becomes "Santana/RS".
options.lowerCaseWordslists the words kept in lower case between two words. The default is the prepositions and the conjunctione:a,ao,aos,à,às,ante,após,até,com,da,das,de,do,dos,e,em,na,nas,no,nos,num,numa,o,para,pela,pelas,pelo,pelos,perante,por,sem,sob,sobre, plus the particlesdel,della,den,der,di,du,vanandvon. The articles are onlyaando. Until 2.4.0ao,aos,à,às,para,pela(s),pelo(s),sob,sobre,até,num,numa,ante,apósandperantewere capitalized.options.upperCaseWordslists the words always in upper case. The default is the company designations (CIA,EIRELI,EPP,LTDA,ME,MEI,S.A,S.A.,S.S.,S/A,S/S,SCP), the document abbreviations (CEP,CNPJ,CPF,RG,UF) and the roman numerals. The match ignores case.S/AandS/Sare matched across the slash even though a slash splits words ("casa de carnes s/a"becomes"Casa de Carnes S/A"). A list given replaces its default, so{ upperCaseWords: [] }turns"empresa ltda"into"Empresa Ltda". A state code after a/, or ending the value after a spaced hyphen, a spaced en dash or a comma, stays upper case even withupperCaseWordsgiven.- Words split at whitespace,
-,/, apostrophes and adjoining punctuation. Runs of whitespace (spaces, tabs, newlines) collapse into one space, and the whitespace at the start and at the end is dropped. - A lower-case word keeps its capital when it is first, last or followed by punctuation.
MEis upper-cased only as a designation (last word, or right before another company designation such asEPPorS/A), never when a hyphen or an apostrophe attaches it to the previous word ("diga-me"becomes"Diga-Me").S.Atyped without the final dot is a designation too ("empresa s.a"becomes"Empresa S.A"; until 2.4.0 it was"Empresa S.a").SAwithout dots is left alone (the surname Sá). - The elided particle
dis lower case wherever it appears, even as the first word, but only when an apostrophe and a word follow it ("santa bárbara d'oeste"becomes"Santa Bárbara d'Oeste";"rua d"becomes"Rua D"). A single letter right after an apostrophe is the English possessive and stays lower case ("bob's"becomes"Bob's"). - Roman numerals from
IItoXXXIXare upper-cased ("rua xxiv de maio"becomes"Rua XXIV de Maio").VIis left out because it is also the verb form "vi", and numerals fromXLon are left out because letters such as L, C and D spell ordinary words. 2.4.0 stopped atXXIII. - A state code (UF) is upper-cased after a
/("santana/rs"becomes"Santana/RS"). As the last word, it is also upper-cased after a spaced hyphen, a spaced en dash or a comma, the "Cidade - UF" form of the Correios addressing guide:"brasília - df"becomes"Brasília - DF". Anywhere else, or after an unspaced hyphen ("brasília-df"), the two letters are an ordinary word. 2.4.0 upper-cased a state code only after/. - Only a real state code counts:
"santana/br"becomes"Santana/Br". - A first letter whose upper case is two letters (
ß) keeps its case:"straße"becomes"Straße"and"ßa"stays"ßa". The letters after the first are lowered one by one, so"İSTANBUL"becomes"İstanbul". - A
lowerCaseWordsorupperCaseWordsthat is not an array falls back to its default, and a member that is not a string is ignored. Avaluethat is not a string returns an empty string.
| Parameter | Type | Required |
|---|---|---|
value | string | yes |
options | CapitalizeOptions | no |
options.lowerCaseWords | string[] | no |
options.upperCaseWords | string[] | no |
| returns | string |
Capitalize the first letter of each word, the way a Brazilian name, company name or address is written, with no options needed.
- Options (
CapitalizeOptions):lowerCaseWords, words kept in lower case between two words, by default the prepositions and the conjunctione, such asde,da,do,ao,para,pelo,sobre,até(the articles are onlyaando);upperCaseWords, words always in upper case, by default company designations and abbreviations such asLTDA,S.A.,ME,CNPJand roman numerals. A list replaces its default. - Words split at whitespace,
-,/, apostrophes and adjoining punctuation; whitespace runs collapse into one space. - A lower-case word that is first, last or followed by punctuation is a designator and keeps its capital.
MEis upper-cased only as a designation (last word, or before another designation);S.Awithout the final dot is a designation too, whileSAwithout dots is left alone (the surname Sá). A state code after a/is upper-cased even withupperCaseWordsgiven, and so is one that ends the value after a spaced-, a spaced–or a,, the Correios' "Cidade – UF".
import { capitalize } from '@brazilian-utils/brazilian-utils';
capitalize('jose da silva'); // Jose da Silva
capitalize('JOSÉ DA SILVA'); // José da Silva
capitalize('empresa ltda'); // Empresa LTDA
capitalize('banco do brasil s.a.'); // Banco do Brasil S.A.
capitalize('casa de carnes s/a'); // Casa de Carnes S/A ("S/A" is matched across the slash)
capitalize('mogi-guaçu'); // Mogi-Guaçu ("-" starts a new word)
capitalize("santa bárbara d'oeste"); // Santa Bárbara d'Oeste ("'" starts a new word, "d" stays lower case)
capitalize("bob's"); // Bob's (a single letter after an apostrophe is the English possessive)
capitalize('rua a, 100'); // Rua A, 100 (a preposition followed by punctuation is a designator)
capitalize('fulano comércio me'); // Fulano Comércio ME ("ME" as the last word is the designation)
capitalize('não-me-toque'); // Não-Me-Toque (anywhere else "me" is an ordinary word)
capitalize('(empresa) ltda'); // (Empresa) LTDA
capitalize('luiz von schmidt'); // Luiz von Schmidt
capitalize('casa para todos'); // Casa para Todos (contracted prepositions such as "ao", "às", "pelo" and "sobre" stay lower case too)
capitalize('empresa s.a'); // Empresa S.A
capitalize('santana/rs'); // Santana/RS ("RS" is a state code right after a "/")
capitalize('porto alegre/rs'); // Porto Alegre/RS
capitalize('brasília - df'); // Brasília - DF (a state code as the last word after " - ", " – " or ", ")
capitalize('santana rs'); // Santana Rs (no "/", so "rs" is just a word)
capitalize('rua xv de novembro'); // Rua XV de Novembro (roman numeral, "de" stays lower case)
capitalize('joão paulo ii'); // João Paulo II
capitalize('rua xxiv de maio'); // Rua XXIV de Maio (roman numerals from II to XXXIX, except VI, the verb "vi")
capitalize('de'); // De (a preposition keeps its capital when it is the first word)
capitalize('empresa ltda', { upperCaseWords: [] }); // Empresa Ltda (the list given replaces the default one)
capitalize('josé Ama MARIA', { lowerCaseWords: ['ama'] }); // José ama Maria
capitalize('doc inválido', { upperCaseWords: ['DOC'] }); // DOC Inválido (case-insensitive match)
capitalize(' josé maria '); // José Maria (every run of whitespace, tabs and newlines included, collapses into one space)Source: Manual de Redação da Presidência da República.
Code: brazilian-utils/javascriptTry it with JavaScript capitalize
Shared test cases (162) and the result in each library text.capitalize
Remove accents
- JavaScript library
- Python library
- Go library
- Ruby library2 cases fail
- Rust library
- .NET library
- Erlang library
Removes diacritical marks (accents, tildes, cedillas) from a string.
- The function decomposes every character (Unicode NFD) and drops every combining mark (general category M), so accents from any script go.
- Letters with no decomposition into a base letter and a mark stay as they are (
ß,ø,æ). Avaluethat is not a string returns an empty string.
| Parameter | Type | Required |
|---|---|---|
value | string | yes |
| returns | string |
Remove diacritical marks (accents, tildes, cedillas) from a string.
- Every combining mark (Unicode general category M) is dropped, so accents from any script go.
import { removeAccents } from '@brazilian-utils/brazilian-utils';
removeAccents('São Paulo'); // 'Sao Paulo'
removeAccents('Piauí'); // 'Piaui'
removeAccents('Ceará'); // 'Ceara'
removeAccents('Açaí'); // 'Acai'
removeAccents(''); // ''Try it with JavaScript removeAccents
Shared test cases (11) and the result in each library text.removeAccents
Official sources
- www4.planalto.gov.br/centrodeestudos/assuntos/…/manual-de-redacao.pdf
- www4.planalto.gov.br/centrodeestudos/assuntos/…/manual-de-redacao-da-presidencia-da-republica
- servicodados.ibge.gov.br/api/docs/…/localidades
- planalto.gov.br/ccivil_03/leis/…/l6404consol.htm
- planalto.gov.br/ccivil_03/leis/…/l10406compilada.htm
- planalto.gov.br/ccivil_03/leis/…/lcp123.htm
- gov.br/empresas-e-negocios/pt-br/…/anexo-iv-limitada_link.pdf
- gov.br/empresas-e-negocios/pt-br/…/instrucoes-normativas
- correios.com.br/enviar/precisa-de-ajuda/…/guia-de-enderecamento
- unicode.org/reports/tr15
- unicode.org/reports/tr44
Last updated on
