Towards a more robust teaching evaluation by articulating M-estimators, the Friedman test, additivity tests, and reliability metrics

Authors

DOI:

https://doi.org/10.56219/dialgica.v23i2.6344

Keywords:

statistical inference, statistical analysis, teacher assessment

Abstract

This research proposes a comprehensive statistical approach to evaluate teaching performance through robust estimators, normality tests, reliability analysis (Cronbach's alpha), Hotelling's contrast and Friedman's test. Using a quantitative methodology with a Likert scale applied to 160 students, the study aims to validate the instrument and identify strengths and areas for improvement. The results show adequate internal consistency and the absence of significant outliers, highlighting organization and feedback as strengths, while the applicability of learning and participative dynamism is identified as weaknesses. The multimethod approach proves to be a robust model, demonstrating that the integration of complementary statistical perspectives offers more solid and nuanced conclusions than any isolated technique.

Author Biography

Melvin Octavio Fiallos Gonzáles, Universidad Pedagógica Nacional Francisco Morazán. Tegucigalpa, D.C. - Honduras

Profesor de educación técnica industrial, Master en Currículo e Ingeniería, con formación en investigación, sistemas de manufactura, Currículo, experiencia docente a nivel superior.

References

Abrami, P.C., d’Apollonia, S., Rosenfield, S. (2007). The dimensionality of student ratings of instruction: What we know and what we do not. En Perry, R.P., Smart, J.C. (eds.) The Scholarship of Teaching and Learning in Higher Education: An Evidence-Based Perspective. Springer. https://doi.org/10.1007/1-4020-5742-3_10

Anderson, T. W. (2003). An introduction to multivariate statistical analysis (2ª ed.). Wiley-Interscience.

Andrews, D. F. (1974). A robust method for multiple linear regression. Technometrics, 16(4), 523-531. https://doi.org/10.1080/00401706.1974.10489233

Andrews, D. F., Bickel, P. J., Hampel, F. R., Huber, P. J., Rogers, W. H., & Tukey, J. W. (1972). Robust estimates of location: Survey and advances. Princeton University Press.

Beaton, A. E., & Tukey, J. W. (1974). The fitting of power series, meaning polynomials, illustrated on band-spectroscopic data. Technometrics, 16(2), 147-185. https://doi.org/10.1080/00401706.1974.10489171

Beltrán, L. A. (2017). Didáctica y Currículo. En Corporación Universitaria Minuto de Dios (Eds.). Didáctica para no didácticos: Reflexiones frente a la didáctica, enseñanzas y experiencias pedagógicas.

Bermejo, B., y Ballesteros, C. (2014). Manual de didáctica general para maestros de educación infantil y de primaria, Pirámide.

Bernal, C. A. (2010). Metodología de la Investigación (3era ed.), Pearson Educación.

Creswell, J. W. (2009). Research Design, qualitative, quantitative, and mixed Methods Approaches, SAGE.

Cook, S., Watson, D., & Webb, R. (2024). Performance evaluation in teaching: Dissecting student evaluations in higher education. Studies in Educational Evaluation, 81, 101342. https://doi.org/10.1016/j.stueduc.2024.101342

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334.

Danielson, C. (2007). Enhancing professional practice: A framework for teaching. Association for Supervision and Curriculum Development.

Fiallos, L., y Fiallos, M. O. (2024). Diseño y validación de un instrumento de investigación desde los constructos cualitativos hasta los cuantitativos. Revista Holón, 2(7), 45–58. https://doi.org/10.48204/j.holon.n7.a6586

Field, A. (2018). Discovering statistics using IBM SPSS statistics. SAGE Publications.

Friedman, M. (1937). The use of ranks to avoid the assumption of normality implicit in the analysis of variance. Journal of the American Statistical Association, 32(200), 675-701.

Friedman, M. (1940). A comparison of alternative tests of significance for the problem of m rankings. The Annals of Mathematical Statistics, 11(1), 86-92.

Frumento, P., & Bottai, M. (2017). Parametric quantile regression: A new approach for robust estimation. Statistical Methods in Medical Research, 26(5), 2083-2099. https://doi.org/10.1007/s10260-021-00557-7

Pineda García, G. L., Banegas Retes, D. L., Fiallos Gonzáles, L., Gómez González, A. Y., & Fiallos Gonzáles, M. O. (2025). Exploratory factor analysis of an instrument that evaluates the relationship between language development theories and methodological processes in third grade of pre-basic education and first grade of basic education. Salud, Ciencia y Tecnología, 5, 2455. https://doi.org/10.56294/saludcyt20252455

George, D., & Mallery, P. (2003). SPSS for Windows step by step: A simple guide and reference, Allyn & Bacon.

Gimeno Sacristán, J. (1998). El curriculum: Una reflexión sobre la práctica. Morata.

Hampel, F. R. (1968). Contributions to the theory of robust estimation [Tesis doctoral no publicada, University of California, Berkeley].

Huber, P. J. (1964). Robust estimation of a location parameter. The Annals of Mathematical Statistics, 35(1), 73-101. https://doi.org/10.1214/aoms/1177703732

Huber, P. J. (1981). Robust statistics. Wiley. https://doi.org/10.1002/0471725250

IBM Corp. (2026). IBM SPSS Statistics: M-estimators subcommand. IBM Documentation. https://www.ibm.com/docs/it/spss-statistics/31.0.0

Maronna, R. A., Martin, R. D., Yohai, V. J., & Salibián-Barrera, M. (2019). Robust statistics: Theory and methods, Wiley.

Cardeña Ojeda, C. A. (2024). Levene’s test for verifying homoscedasticity between groups in quasi-experiments in social sciences. South Eastern European Journal of Public Health, 2119–2125. https://doi.org/10.70135/seejph.vi.2342

Olivos, T. M. (2016). Evaluación del aprendizaje y para el aprendizaje: Reinventar la evaluación en el aula. Universidad Autónoma Metropolitana. https://casadelibrosabiertos.uam.mx/contenido/contenido/Libroelectronico/Evaluacion_del_aprendizaje_.pdf

Organization for Economic Cooperation and Development. (2013). Teachers for the 21st century: Using evaluation to improve teaching, OECD Publishing. https://doi.org/10.1787/9789264193864-en

Perlman, M. D. (2019). On the feasibility of parsimonious variable selection for Hotelling's T2-test. arXiv. https://doi.org/10.48550/arXiv.1910.03669

Pulido, S., y Rodríguez, J. (2014). Estadística descriptiva y análisis cualitativo. Universidad Nacional de Colombia.

Quansah, F., Asamoah, D., Amankwah, B., & Agormedah, E. K. (2024). Validity of student evaluation of teaching in higher education: A systematic review. Frontiers in Education, 9, 1329734. https://doi.org/10.3389/feduc.2024.1329734

Razali, N. M., & Wah, Y. B. (2011). Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling tests. Journal of Statistical Modeling and Analytics, 2(1), 21-33. https://www.nrc.gov/docs/ML1714/ML17143A100.pdf

Rhemtulla, M., Brosseau-Liard, P. É., & Savalei, V. (2012). When can categorical variables be treated as continuous? A comparison of robust continuous and categorical SEM estimation methods under suboptimal conditions. Psychological Methods, 17(3), 354-373. https://doi.apa.org/doi/10.1037/a0029315

Sepúlveda García, J. M., Suárez-Giraldo, A. M., Rodas-Rodríguez, J. M., Ruiz-Ortega, F. J., y Henao-Henao, M. D. (2018). Validación y aplicación de un test modificado de Vandenberg y Kuse de rotación mental para simetría molecular. Tecné, Episteme y Didaxis: TED, (43), 155–171. https://doi.org/10.17227/ted.num43-8656

Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870

Tukey, J. W. (1960). A survey of sampling from contaminated distributions. En I. Olkin, Contributions to probability and statistics: Essays in honor of Harold Hotelling. (448-485). Stanford University Press.

Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007

Wilcox, R. R. (2017). Introduction to robust estimation and hypothesis testing, Academic Press.

Zabaleta, F. (2007). The use and misuse of student evaluations of teaching. Teaching in Higher Education, 12(1), 55–76. https://doi.org/10.1080/1356251060110213

Published

2026-07-20

How to Cite

Fiallos Gonzáles, M. O. (2026). Towards a more robust teaching evaluation by articulating M-estimators, the Friedman test, additivity tests, and reliability metrics. DIALOGICA, 23(2), 747–769. https://doi.org/10.56219/dialgica.v23i2.6344