Extracting headings' text from word doc

Tags:

I am trying to extract text from headings(of any level) in a MS Word document(.docx file). Currently I am trying to solve using python-docx, but unfortunately I am still not able to figure out if it is even feasible after reading it(maybe I am mistaken).

I tried to look for the solutions online but found nothing specific to my task. It would be great if someone could guide me here.

806

asked Nov 02 '16 20:11

Muhammad Muaz

1 Answers

The fundamental challenge is identifying heading paragraphs. There's nothing stopping an author from formatting a "regular" paragraph to look like (and serve as) a heading as far as a reader is concerned.

However, it's not uncommon for authors to reliably use styles to create headings, because doing so makes it possible to automatically compile those headings into a table of contents.

In that case, you can just iterate over the paragraphs, and pick out those with one of the heading styles.

def iter_headings(paragraphs):
    for paragraph in paragraphs:
        if paragraph.style.name.startswith('Heading'):
            yield paragraph

for heading in iter_headings(document.paragraphs):
    print heading.text

Heading levels may be parsed from the full style name if they've kept the defaults (like 'Heading 1', 'Heading 2', ...).

This may need to be adjusted if the author has renamed the heading styles.

There are more sophisticated approaches which are more reliable (as far as being style-name independent), but those don't have API support so you'd need to dig into the internal code and interact with some of the style XML directly I expect.

145

answered Sep 22 '22 12:09

scanny

Related questions
                            
                                Map to List error: Series object not callable
                            
                                class attribute lookup rule?
                            
                                Flask admin overrides password when user model is changed
                            
                                Numpy Vectorization of sliding-window operation
                            
                                How to sum and to mean one DataFrame to create another DataFrame
                            
                                Mock property return value gets overridden when instantiating mock object
                            
                                Matplotlib figure size in Jupyter reset by inlining in Jupyter
                            
                                Multiprocessing Pool - how to cancel all running processes if one returns the desired result?
                            
                                Compute the running (cumulative) maximum for a series in pandas
                            
                                how to initialize multiple columns to existing pandas DataFrame
                            
                                plot 2d lines by line equation in Python using Matplotlib
                            
                                Python, error with web driver (Selenium)
                            
                                Matplotlib "pick_event" not working in embedded graph with FigureCanvasTkAgg
                            
                                vlookup between 2 Pandas dataframes
                            
                                Python:Update list of tuples
                            
                                Replace 0 with blank in dataframe Python pandas
                            
                                Why are many Python built-in/standard library functions actually classes
                            
                                Pycharm does not recognize Cython modules located in path
                            
                                Is there anyway Google App Engine apps can communicate or control Machine Learning models or tasks?
                            
                                PonyORM: What is the most efficient way to add new items to a pony database without knowing which items already exist?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Extracting headings' text from word doc

Tags:

python

text

parsing

ms-word

python-docx

Muhammad Muaz

People also ask

1 Answers

scanny

Recent Activity

Donate For Us