光学字符识别 (OCR)
Vision API 可以检测并提取图片中的文本。支持光学字符识别 (OCR) 的注释功能有两种:
TEXT_DETECTION可检测并提取任何图片中的文本。例如,某张照片可能包含街道标志或交通标志。JSON 包含所提取的整个字符串,以及各个字词及其边界框。
DOCUMENT_TEXT_DETECTION也可提取图片中的文本,但其响应针对密集文本和文档进行了优化。JSON 包含页面、文本块、段落、字词和换行信息。
亲自尝试
如果您是 Google Cloud 新手,请创建一个账号来评估 Cloud Vision 在实际场景中的表现。新客户还可获享 $300 赠金,用于运行、测试和部署工作负载。
免费试用 Cloud Vision文本检测请求
设置您的 Google Cloud 项目和身份验证
如果您尚未创建 Google Cloud 项目,请立即创建。展开本部分可查看相关说明。
- Sign in to your Google Cloud account. If you're new to Google Cloud, create an account to evaluate how our products perform in real-world scenarios. New customers also get $300 in free credits to run, test, and deploy workloads.
-
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
-
Enable the Vision API.
Roles required to enable APIs
To enable APIs, you need the Service Usage Admin IAM role (
roles/serviceusage.serviceUsageAdmin), which contains theserviceusage.services.enablepermission. Learn how to grant roles. -
Install the Google Cloud CLI.
-
如果您使用的是外部身份提供方 (IdP),则必须先使用联合身份登录 gcloud CLI。
-
如需初始化 gcloud CLI,请运行以下命令:
gcloud init -
In the Google Cloud console, on the project selector page, select or create a Google Cloud project.
Roles required to select or create a project
- Select a project: Selecting a project doesn't require a specific IAM role—you can select any project that you've been granted a role on.
-
Create a project: To create a project, you need the Project Creator role
(
roles/resourcemanager.projectCreator), which contains theresourcemanager.projects.createpermission. Learn how to grant roles.
-
Verify that billing is enabled for your Google Cloud project.
-
Enable the Vision API.
Roles required to enable APIs
To enable APIs, you need the Service Usage Admin IAM role (
roles/serviceusage.serviceUsageAdmin), which contains theserviceusage.services.enablepermission. Learn how to grant roles. -
Install the Google Cloud CLI.
-
如果您使用的是外部身份提供方 (IdP),则必须先使用联合身份登录 gcloud CLI。
-
如需初始化 gcloud CLI,请运行以下命令:
gcloud init - BASE64_ENCODED_IMAGE:二进制图片数据的 base64 表示(ASCII 字符串)。此字符串应类似于以下字符串:
/9j/4QAYRXhpZgAA...9tAVx/zDQDlGxn//2Q==
- PROJECT_ID:您的 Google Cloud 项目 ID。
检测本地图片中的文本
您可以使用 Vision API 对本地图片文件执行特征检测。
对于 REST 请求,请将图片文件的内容作为 base64 编码的字符串在请求正文中发送。
对于 gcloud 和客户端库请求,请在请求中指定本地图片的路径。
gcloud
如需执行文本检测,请使用 gcloud ml vision detect-text 命令,如以下示例所示:
gcloud ml vision detect-text ./path/to/local/file.jpg
REST
在使用任何请求数据之前,请先进行以下替换:
HTTP 方法和网址:
POST https://vision.googleapis.com/v1/images:annotate
请求 JSON 正文:
{
"requests": [
{
"image": {
"content": "BASE64_ENCODED_IMAGE"
},
"features": [
{
"type": "TEXT_DETECTION"
}
]
}
]
}
如需发送请求,请选择以下方式之一:
curl
将请求正文保存在名为 request.json 的文件中,然后执行以下命令:
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "x-goog-user-project: PROJECT_ID" \
-H "Content-Type: application/json; charset=utf-8" \
-d @request.json \
"https://vision.googleapis.com/v1/images:annotate"
PowerShell
将请求正文保存在名为 request.json 的文件中,然后执行以下命令:
$cred = gcloud auth print-access-token
$headers = @{ "Authorization" = "Bearer $cred"; "x-goog-user-project" = "PROJECT_ID" }
Invoke-WebRequest `
-Method POST `
-Headers $headers `
-ContentType: "application/json; charset=utf-8" `
-InFile request.json `
-Uri "https://vision.googleapis.com/v1/images:annotate" | Select-Object -Expand Content
如果请求成功,服务器将返回一个 200 OK HTTP 状态代码以及 JSON 格式的响应。
TEXT_DETECTION 响应包含检测到的词组及其边界框,以及各个字词及其边界框。
响应
{
"responses": [
{
"textAnnotations": [
{
"locale": "en",
"description": "WAITING?\nPLEASE\nTURN OFF\nYOUR\nENGINE\n",
"boundingPoly": {
"vertices": [
{
"x": 341,
"y": 828
},
{
"x": 2249,
"y": 828
},
{
"x": 2249,
"y": 1993
},
{
"x": 341,
"y": 1993
}
]
}
},
{
"description": "WAITING?",
"boundingPoly": {
"vertices": [
{
"x": 352,
"y": 828
},
{
"x": 2248,
"y": 911
},
{
"x": 2238,
"y": 1148
},
{
"x": 342,
"y": 1065
}
]
}
},
{
"description": "PLEASE",
"boundingPoly": {
"vertices": [
{
"x": 1210,
"y": 1233
},
{
"x": 1907,
"y": 1263
},
{
"x": 1902,
"y": 1383
},
{
"x": 1205,
"y": 1353
}
]
}
},
{
"description": "TURN",
"boundingPoly": {
"vertices": [
{
"x": 1210,
"y": 1418
},
{
"x": 1730,
"y": 1441
},
{
"x": 1724,
"y": 1564
},
{
"x": 1205,
"y": 1541
}
]
}
},
{
"description": "OFF",
"boundingPoly": {
"vertices": [
{
"x": 1792,
"y": 1443
},
{
"x": 2128,
"y": 1458
},
{
"x": 2122,
"y": 1581
},
{
"x": 1787,
"y": 1566
}
]
}
},
{
"description": "YOUR",
"boundingPoly": {
"vertices": [
{
"x": 1219,
"y": 1603
},
{
"x": 1746,
"y": 1629
},
{
"x": 1740,
"y": 1759
},
{
"x": 1213,
"y": 1733
}
]
}
},
{
"description": "ENGINE",
"boundingPoly": {
"vertices": [
{
"x": 1222,
"y": 1771
},
{
"x": 1944,
"y": 1834
},
{
"x": 1930,
"y": 1992
},
{
"x": 1208,
"y": 1928
}
]
}
}
],
"fullTextAnnotation": {
"pages": [
...
]
},
"paragraphs": [
...
]
},
"words": [
...
},
"symbols": [
...
}
]
}
],
"blockType": "TEXT"
},
...
]
}
],
"text": "WAITING?\nPLEASE\nTURN OFF\nYOUR\nENGINE\n"
}
}
]
}
Go
试用此示例之前,请按照《Vision 快速入门:使用客户端库》中的 Go 设置说明进行操作。 如需了解详情,请参阅 Vision Go API 参考文档。
如需向 Vision 进行身份验证,请设置应用默认凭证。如需了解详情,请参阅为本地开发环境设置身份验证。
// detectText gets text from the Vision API for an image at the given file path.
func detectText(w io.Writer, file string) error {
ctx := context.Background()
client, err := vision.NewImageAnnotatorClient(ctx)
if err != nil {
return err
}
f, err := os.Open(file)
if err != nil {
return err
}
defer f.Close()
image, err := vision.NewImageFromReader(f)
if err != nil {
return err
}
annotations, err := client.DetectTexts(ctx, image, nil, 10)
if err != nil {
return err
}
if len(annotations) == 0 {
fmt.Fprintln(w, "No text found.")
} else {
fmt.Fprintln(w, "Text:")
for _, annotation := range annotations {
fmt.Fprintf(w, "%q\n", annotation.Description)
}
}
return nil
}
Java
在试用此示例之前,请按照Vision API 快速入门:使用客户端库中的 Java 设置说明进行操作。如需了解详情,请参阅 Vision API Java 参考文档。
import com.google.cloud.vision.v1.AnnotateImageRequest;
import com.google.cloud.vision.v1.AnnotateImageResponse;
import com.google.cloud.vision.v1.BatchAnnotateImagesResponse;
import com.google.cloud.vision.v1.EntityAnnotation;
import com.google.cloud.vision.v1.Feature;
import com.google.cloud.vision.v1.Image;
import com.google.cloud.vision.v1.ImageAnnotatorClient;
import com.google.protobuf.ByteString;
import java.io.FileInputStream;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;
public class DetectText {
public static void