Search CORE

40 research outputs found

A practical approach to language complexity: a wikipedia case study

Author: A Halavais
A Kornai
A Mikheev
András Kornai
D van Leijenhorst
D Varga
E Gabrilovich
Eduardo G. Altmann
EG Altmann
EG Altmann
F Tweedie
GR Klare
JC Roberts
János Kertész
M Serrano
MD Besten
MK Paasche-Orlow
O Medelyan
R Baeza Yates
R Gunning
R Lambiotte
S Javanmardi
T Yasseri
T Yasseri
Taha Yasseri
Publication venue: 'Public Library of Science (PLoS)'
Publication date: 01/01/2012
Field of study

In this paper we present statistical analysis of English texts from Wikipedia. We try to address the issue of language complexity empirically by comparing the simple English Wikipedia (Simple) to comparable samples of the main English Wikipedia (Main). Simple is supposed to use a more simplified language with a limited vocabulary, and editors are explicitly requested to follow this guideline, yet in practice the vocabulary richness of both samples are at the same level. Detailed analysis of longer units (n-grams of words and part of speech tags) shows that the language of Simple is less complex than that of Main primarily due to the use of shorter sentences, as opposed to drastically simplified syntax or vocabulary. Comparing the two language varieties by the Gunning readability index supports this conclusion. We also report on the topical dependence of language complexity, that is, that the language is more advanced in conceptual articles compared to person-based (biographical) and object-based articles. Finally, we investigate the relation between conflict and language complexity by analyzing the content of the talk pages associated to controversial and peacefully developing articles, concluding that controversy has the effect of reducing language complexity

arXiv.org e-Print Archive

Crossref

SZTAKI Publication Repository

Directory of Open Access Journals

PubMed Central

The Francis Crick Institute

ReadNet: A Hierarchical Transformer Framework for Web Article Readability Analysis

Author: A Coxhead
AC Graesser
DS McNamara
DS McNamara
E Dale
E Fry
E Gibson
EB Fry
GH Mc Laughlin
GR Klare
J Anderson
J Duchi
K Collins-Thompson
K Collins-Thompson
M Chen
M Chen
M Coleman
M Louwerse
O De Clercq
R Gunning
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 06/03/2021
Field of study

Analyzing the readability of articles has been an important sociolinguistic task. Addressing this task is necessary to the automatic recommendation of appropriate articles to readers with different comprehension abilities, and it further benefits education systems, web information systems, and digital libraries. Current methods for assessing readability employ empirical measures or statistical learning techniques that are limited by their ability to characterize complex patterns such as article structures and semantic meanings of sentences. In this paper, we propose a new and comprehensive framework which uses a hierarchical self-attention model to analyze document readability. In this model, measurements of sentence-level difficulty are captured along with the semantic meanings of each sentence. Additionally, the sentence-level features are incorporated to characterize the overall readability of an article with consideration of article structures. We evaluate our proposed approach on three widely-used benchmark datasets against several strong baseline approaches. Experimental results show that our proposed method achieves the state-of-the-art performance on estimating the readability for various web articles and literature.Comment: ECIR 202

arXiv.org e-Print Archive

Crossref

Readability of medicinal package leaflets: a systematic review

Author: Afonso Cavaco
Brosnan S
Calamusa A
Carla Pires
Carrigan N
Cavaco A
Cavaco A
Cavaco AM
Dale E
Dowse R
DuBay WH
Flesch R
Franck MCJ
Fuchs J
Fuchs J
Gazmarariana JA
Gunning R
Klare GR
Knapp P
Knapp P
Lee IH
Leiderman DB
Maat HP
March Cerdá JC
Marina Vigário
Pinero-Lopez MA
Roskos SE
Shiffman S
Symonds T
Wallace LS
Weiss SM
Wolf A
Zite NB
Publication venue: 'FapUNIFESP (SciELO)'
Publication date: 01/01/2015
Field of study

Crossref

Kinetics of Rapid Covalent Bond Formation of Aniline with Humic Acid: ESR Investigations with Nitroxide Spin Labels

Author: A Gulkowska
A Gulkowska
A Jezierski
AD Steen
AP Todd
B Gevao
C Achtnicht
C Lattao
D Colón
E Barriuso
EJ Weber
EJ Weber
F Führ
G Dawel
GE Parris
GR Buettner
H Li
Heinz-Jürgen Steinhoff
HM Bialk
HM Bialk
HM Bialk
J Fuchs
J-M Bollag
J-M Bollaq
JP Klare
K Stolze
KA Thorn
KA Thorn
Kalman Hideg
Kevin Glinka
L Urban
M Förster
M Kawahigashi
M Kästner
Marius Theiling
Michael Matthies
N Senesi
OECD Guideline for the Testing of Chemicals no. 308
P Franchi
S Gadányi
STJ Droge
T Müller
Publication venue: 'Springer Science and Business Media LLC'
Publication date
Field of study

Crossref

Readability Formula for Chinese as a Second Language

Author: GR Klare
Publication venue: 'Springer Science and Business Media LLC'
Publication date
Field of study

Crossref

Readability of consent forms in veterinary clinical research

Author: Chall J. S.
Flesch R
Grabeel KL
Klare GR
Kutner M
McLaughlin GH
Pandiya A
Weiss BD
Publication venue: 'Wiley'
Publication date
Field of study

Crossref

Predicting Text Readability with Personal Pronouns

Author: CS Meppelink
D Biber
E Dale
GHM Laughlin
GR Klare
JE Brinton
K Kotani
R Flesch
R Gunning
S Botas
WH Dubay
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 02/11/2018
Field of study

Part 5: Perceptual IntelligenceInternational audienceWhile the classic Readability Formula exploits word and sentence length, we aim to test whether Personal Pronouns (PPs) can be used to predict text readability with similar accuracy or not. Out of this motivation, we first calculated readability score of randomly selected texts of nine genres from the British National Corpus (BNC). Then we used Multiple Linear Regression (MLR) to determine the degree to which readability could be explained by any of the 38 individual or combinational subsets of various PPs in their orthographical forms (including I, me, we, us, you, he, him, she, her (the Objective Case), it, they and them). Results show that (1) subsets of plural PPs can be more predicative than those of singular ones; (2) subsets of Objective forms can make better predictions than those of Subjective ones; (3) both the subsets of first- and third-person PPs show stronger predictive power than those of second-person PPs; (4) adding the article the to the subsets could only improve the prediction slightly. Reevaluation with resampled texts from BNC verify the practicality of using PPs as an alternative approach to predict text readability

Crossref

Influence of Term Familiarity in Readability of Spanish e-Government Web Information

Author: A Keselman
A Koohang
C Schaffer
D Kauchak
E Dale
F Cuetos
G Leroy
GR Klare
IH Witten
J Fernández-Huerta
J Morato
M Nomura
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 31/10/2020
Field of study

It is well known that linguistic features of a written text affect its readability, understanding readability as the ease with which a reader can understand the text. This paper is focused on the analysing of the influence of some linguistic features on the readability of current Spanish e-Government websites. Specifically, the “familiarity” of the terms on web pages, as well as the “frequency” of these terms are studied, among others. Firstly, this research has analysed a corpus extracted from the current information websites of the Spanish eGovernment and its simplified counterparts. Then, using machine learning methods, a supervised model is built on the influence of different term familiarity lists on text readability in the corpus. Different term lists have been tested and it has been concluded that the differences between them have a great impact on their performance. An accuracy of 81% has been achieved with a combination of frequency lists. As a conclusion, term lists and the frequencies of the terms allow to determine to a high degree the difficulty of understanding the text.Work supported by the Spanish Ministry of Economy, Industry and Competitiveness, (CSO2017-86747-R)

Crossref

e-Archivo (Univ. Carlos III de Madrid e-Archivo)

Using the Readability Assessment Instrument to Evaluate Patient Medication Leaflets

Author: AJ Moe
American Pharmaceutical Association.
CC Doak
GR Klare
J Singh
JE Readence
JF Baumann
MAK Halliday
RM Schulz
Russell J. Sojourner
TH Anderson
TL Harris
United States Pharmacopeial Convention.
Publication venue: 'SAGE Publications'
Publication date
Field of study

Crossref

Briefings

Author: Abrahamsen R.
Adams gR. H.
Ake C.
Ansoms A.
Ayandele E.
Best S.
Cummins I.
Diamond L.
Dollar D.
Klare M.
Obi C.
Okonta I.
Olusanya G.
Peel M.
Ravallion M.
Ravallion M.
Valle V.
Publication venue: 'Informa UK Limited'
Publication date
Field of study

Crossref