Researchers from Carnegie Mellon University created an automated computer system that makes online communications more polite. The system takes non-polite directives and requests and restructures them or adds words to create well-mannered sentences.
Researchers at Carnegie Mellon University have developed an automated method for making communications more polite. This image shows one possible implementation of this method: automated suggestions for improving politeness in emails. Source: Carnegie Mellon University
The team created a dataset of 1.39 million sentences labeled for politeness. These sentences were gathered from emails between employees of Enron until its demise in 2001. The Texas-based energy company was known for its corporate fraud and corruption. Half a million corporate emails become public thanks to the lawsuits surrounding the fraud scandal.
After gathering the sentences to train the system, the team had to define politeness. This was challenging because politeness varies across cultures. What may be deemed as polite in one culture could be an insult in another. The researchers decided to restrict their work to speakers of North American English in a formal setting.
The politeness dataset was analyzed to determine the frequency and distribution of works in polite and nonpolite sentences. The team created a tag to generate a pipeline and perform politeness transfers. The system tags impolite or nonpolite words or phrases and a text generator replaces the tagged item with a polite phrase. The goal is not to change the meaning of the sentence, just adjusting it to be polite.
As the system was completing training, the new sentences became more realistic and could implement subtle changes, like changing first-person singular pronouns were replaced by first-person plural pronouns. Also, instead of just adding "please" to the beginning of the sentence, it integrated "please" into the middle of the sentence so it seemed more natural.
The team is presenting this research at the Association for Computational Linguistics Annual Meeting on July 5, 2020.
