使用watsonx Code Assistant使用本地IBM Granite模型的个人
watsonx Code Assistant个人
对于个人用户,watsonx Code Assistant 可以通过 Ollama 访问本地模型,Ollama 是一种广泛使用的大型语言模型本地推理引擎。 Ollama 封装了底层模型服务项目 llama.cpp。
要提高性能并为您的组织提供全套功能,请在IBM Cloud 上提供watsonx Code Assistant的试用版。 有关详细信息,请参阅 在IBM Cloud中设置watsonx Code Assistant服务。
安装 watsonx Code Assistant 扩展
您可以设置 Ollama,以便在 MicrosoftVisual Studio Code 中使用。
Eclipse IDE 插件不支持此设置。 仅适用于 Visual Studio Code 扩展名。
-
在Visual Studio Marketplace中打开该页面。watsonx Code Assistant Visual Studio Marketplace页面。
-
单击市场页面上的安装。
-
在Visual Studio Code 中,点击安装扩展。
-
在扩展设置中,将 Wca:后端提供程序设置为 ollama。
IBM 员工请注意:如果您是使用 WCA@IBM 内部扩展名的 IBM 或 Red Hat 员工,更改后端提供商将始终恢复为
wcaCore。 更改后台提供商:- 在 WCA@IBM 扩展设置中,清除为 watsonx Code Assistant 启用 WCA@IBM 模式
- 作为替代方案,禁用 WCA@IBM 扩展名
安装奥拉玛
-
下载并运行 ollama 安装程序。
-
在 macOS, 上,您还可以使用 Homebrew 安装 Ollama:
brew install ollama
启动奥拉玛推理服务器
在控制台窗口中运行
ollama serve
使用 Ollama 时,请打开该窗口。
如果收到信息 "Error: listen tcp 127.0.0.1:11434: bind: address already in use,则 Ollama 服务器已经启动。
安装IBM Granite代码模型
首先安装 Ollama 库中 的 "granite-code:8b 模型。
-
打开一个新的控制台窗口。
-
在命令行中输入 "
ollama run granite-code:8b,下载并部署模型。 您会看到类似于以下示例的输出:pulling manifest pulling 8718ec280572... 100% ▕███████████████████████ 4.6 GB pulling e50df8490144... 100% ▕███████████████████████ ▏ 123 B pulling 58d1e17ffe51... 100% ▕███████████████████████▏ 11 KB pulling 9893bb2c2917... 100% ▕███████████████████████▏ 108 B pulling 0e851433eda0... 100% ▕███████████████████████▏ 485 B verifying sha256 digest writing manifest removing any unused layers success >>> -
Type
/byeafter the>>>to exit the Ollama command shell. -
通过键入试试模型:
ollama run granite-code:8b "How do I create a python class?" -
您应该会看到类似的回复:
To create a Python class, you can define a new class using the "class" keyword followed by the name of the class and a colon. Inside the class definition, you can specify the methods and attributes that the class will have. Here is an example: ...
配置 Ollama 主机
默认情况下,Ollama服务器运行在IP地址 127.0.0.1、端口 11434 和http协议上。 如果更改了 Ollama 的 IP 地址或端口:
-
在Visual Studio Code中,打开watsonx Code Assistant 的扩展设置。
-
在 Wca > Local:API 主机,添加主机 IP 和端口。
配置使用的Granite模型
默认情况下,watsonx Code Assistant会为聊天和代码自动补全使用 "granite-code:8b 模型。 如果您的环境有足够的容量,请安装 "granite-code:8b-base 型号。
使用不同的模式:
-
安装 "
granite-code:8b-base模型。 请参阅 安装IBM Granite代码模型。 -
在Visual Studio Code中,打开watsonx Code Assistant 的扩展设置。
-
在 Wca > Local:代码生成模型中,输入 "
granite-code:8b-base。
确保设置安全
默认情况下,Ollama 服务器运行在本地设备的 IP 地址127.0.0.1 上,端口 11434,使用 http 作为协议。 要使用https或通过代理服务器访问,请参阅 Ollama文档。
从本地模型切换到IBM Cloud
您可能决定从本地模型切换到使用IBM Cloud 上的服务实例。 然后,您可以配置Visual Studio Code从本地模型切换到IBM Cloud。
有关详细信息,请参阅 在IBM Cloud中设置watsonx Code Assistant服务。
更新Visual Studio Code编辑器,使用IBM Cloud代替 Ollama:
-
退出 Visual Studio Code。
-
退出 Ollama 应用程序。
-
启动Visual Studio Code,然后打开watsonx Code Assistant。 您应该会看到 "
Ollama is not running in your IDE.信息 -
单击
Switch to watsonx Code Assistant on IBM Cloud。
如果想使用其他方法,可以更改扩展设置:
-
在Visual Studio Code中,打开watsonx Code Assistant 的扩展设置。
-
在 Wca: Backend Provider 中,将“
ollama切换为”wcaCore。 -
重新启动扩展程序以应用更改。