Experiment No. -14 :- Scanner - Installation, Configuration, Using Automatic Document Feeder (ADF), and Optical Character Recognition (OCR)

SCANNER ARCHITECTURE & OCR WORKFLOW DIAGRAM 1. Hardware Mechanisms (Flatbed & ADF) Glass Platen / CCD Optical Carriage ADF Paper Input Tray Flatbed: Single Page / Books / ID Cards ADF: Continuous Multi-Page Batch Feed 2. OCR Extraction Software Pipeline Scanned Image (.BMP / .JPG) OCR Processing Engine Matrix Matching & Feature Analysis Editable Digital Text (.DOCX / .TXT / Searchable PDF) Searchable & Selectable Text HARDWARE FEEDING: Flatbed (glass) vs. ADF (multi-page tray) | SOFTWARE CONVERSION: OCR converts pixel bitmaps into editable text characters.
Experiment No. 14

Scanner - Installation, Configuration, Using Automatic Document Feeder (ADF), and Optical Character Recognition (OCR)

1. Aim of the Experiment

1. To perform physical installation, interface connection, and driver configuration (TWAIN/WIA) for a desktop Scanner on a Windows PC.
2. To execute single-page flatbed scanning and high-speed batch scanning using the Automatic Document Feeder (ADF).
3. To process scanned bitmap document images using Optical Character Recognition (OCR) software to convert image text into editable Word (.docx), Text (.txt), or searchable PDF formats.
OPTIMAL SCANNER DPI CONFIGURATION RULE:
Setting resolution to 300 DPI (Dots Per Inch) in Grayscale or Black & White mode provides the ideal balance for Optical Character Recognition (OCR). Scanning text documents at 600+ DPI creates excessively large file sizes without improving text extraction accuracy, while scanning below 200 DPI causes pixelation that breaks character recognition algorithms.

2. Tools & Software Requirements

  • Hardware: Personal Computer (PC) running Windows OS, Flatbed Scanner with integrated Automatic Document Feeder (ADF) tray or Multifunction Printer (MFP) scanner unit.
  • Cables & Test Media: USB 2.0/3.0 A-to-B Interface Cable, Power Supply Adapter, Single A4 document sheets, Multi-page document stack (5-10 pages), Printed text sample sheets for OCR.
  • Software & Drivers: TWAIN / WIA Scanner Drivers, OEM Scan Utility (e.g., HP Smart, Epson Scan, Canon IJ Scan Utility, or Windows Fax & Scan), OCR Software Engine (e.g., ABBYY FineReader, Readiris, FreeOCR, or OneNote/Google Docs OCR).

3. Step-by-Step Procedure

PART A: Scanner Hardware Installation & Driver Configuration

  1. Unpack the scanner unit and place it on a stable, level surface. Locate the optical transport lock switch at the bottom/side of the flatbed unit and slide it to the Unlocked position.
  2. Connect the USB interface cable between the scanner and an available USB port on the PC. Plug in the power supply adapter and switch on the unit.
  3. Insert the driver disk or download the latest TWAIN/WIA driver software package from the manufacturer's website. Run the setup file as Administrator.
  4. Verify installation: Open Windows Control Panel → Devices and Printers → Scanners and Cameras. Ensure the scanner model is listed with an "Online / Ready" status. Perform a test scan to verify driver communication.

PART B: Operating Flatbed vs. Automatic Document Feeder (ADF)

  1. Flatbed Scan Mode (Single Page / ID Cards): Lift the scanner top lid. Align the document face-down on the platen glass against the top-left corner indicator arrow. Close the lid. Launch the scan software, select Document Source: Flatbed Glass, set resolution to 300 DPI, and click Scan.
  2. ADF Batch Scan Mode (Multi-Page Documents): Adjust the paper guides on the top ADF feed tray. Neat-stack a 5-to-10-page document batch and slide it face-up into the ADF input tray until the paper sensor engages (indicated by an audible beep or LED).
  3. In the scan software interface, change Document Source to Feeder / ADF. Configure page orientation, set scan mode to Duplex (2-Sided) or Simplex (1-Sided), select PDF format, and click Scan.
  4. Observe the motorized pickup rollers feeding each sheet continuously into the output tray to generate a unified multi-page PDF document.

PART C: Optical Character Recognition (OCR) Processing

  1. Launch your designated OCR application (e.g., ABBYY FineReader, Readiris, FreeOCR, or cloud OCR via Google Docs / OneNote).
  2. Import the scanned image file (.JPG, .BMP, or non-searchable .PDF) created in Part B.
  3. Set the primary document language (e.g., English) and select the recognition mode (e.g., Extract Plain Text or Maintain Page Layout & Formatting).
  4. Click the Process / Perform OCR button. The software analyzes the light and dark pixel patterns, comparing them against internal font matrices to extract characters.
  5. Compare the original scanned image against the extracted text panel in the software window. Correct any unrecognized characters or typos.
  6. Export the final converted document as an editable Microsoft Word document (.docx) or Searchable PDF (.pdf).

4. Technical Comparison & Observation Table

Scanning Parameter Flatbed Scanning Mode Automatic Document Feeder (ADF)
Primary Use Case Thick books, passports, ID cards, fragile single sheets Multi-page office reports, contracts, double-sided batch forms
Media Handling Mechanism Stationary document placed manually on glass platen Motorized roller feed pulling stack past fixed sensor
Operation Speed Slow (requires manual sheet swapping per page) High-Speed (Continuous automated multi-page feed)
Optimal DPI for OCR 300 DPI (Grayscale or Black & White) 300 DPI (1-Bit Monochrome or Grayscale)
Risk / Maintenance Factor Fingerprints and dust on glass surface Paper jams, staple damage, multi-sheet pickup wear

5. Viva-Voce / Oral Examination Preparation

Q1. What do the terms TWAIN and WIA stand for in scanner driver configuration?

Ans: TWAIN (Technology Without An Interesting Name) and WIA (Windows Image Acquisition) are standard application programming interfaces (APIs) and driver protocols that allow software applications (e.g., Photoshop, Word) to communicate directly with scanner hardware.

Q2. How does an Optical Character Recognition (OCR) software engine convert images into text?

Ans: OCR software inspects the light and dark pixels of a scanned bitmap image. It isolates character shapes through matrix matching (comparing shapes to stored font models) and feature extraction (analyzing lines, loops, and intersections) to translate image pixels into editable ASCII/Unicode text codes.

Q3. What is the key functional difference between a standard PDF and a Searchable PDF?

Ans: A standard scanned PDF is simply a container holding a flat raster image of the page. A Searchable PDF contains the original page image with an invisible layer of text placed directly beneath it, allowing users to search, highlight, and copy text seamlessly.

6. Practical Result & Conclusion

Result: Scanner interface hardware installation and driver installation (TWAIN/WIA) were completed successfully. Single-page flatbed scanning and high-speed multi-page batch scanning using the Automatic Document Feeder (ADF) were performed. Scanned bitmap images were processed through OCR software and successfully exported into editable text formats.

Post a Comment

0 Comments